Good Start Labs trains AI on games for real-world tasks
Good Start Labs has demonstrated that training AI models on strategy games like Diplomacy and 1830 can successfully transfer reasoning and execution skills to real-world business tasks.

Good Start Labs, which launched with 3.6 million dollars in backing from Inovia, General Catalyst, Every, and angel investors, is demonstrating how virtual gaming environments can prepare AI for professional workflows.
Co-founders Alex Duffy and Tyler Marques recently trained a 30B model inside the strategy game 1830: The Game of Railroads and Robber Barons. When evaluating the model on financial research tasks, the team discovered that only a multi-turn terminal agent design—which uses tools to plan and adapt—successfully improved performance on the Finance-Agent benchmark. A simpler single-turn question-answering design failed to transfer those skills.
This is not the first time the company has linked gaming to real-world capabilities. Duffy noted that fine-tuning models on the strategy game Diplomacy improved their scores on customer support and industrial operations benchmarks. The startup has analyzed how various frontier models behave in these environments. For instance, OpenAI's o3 model won Diplomacy matches by planning future betrayals, while Claude Opus 4 refused to lie and lost. Good Start Labs also tracks newer models like Grok 4 Fast, Gemini 2.5 Pro, Claude Fable 5.1, and GPT-6 Astra. On the LOL-Alignment benchmark, which uses over 14,000 hands from the game playbadcards, Astra topped the rankings by matching human judges 50 percent of the time, just two percent ahead of GPT-4o.
For AI practitioners, these findings suggest that the design of the training harness is critical for teaching specific capabilities. In their co-authored paper, COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks, the founders outline how a decision agent can access a learnable skill bank to guide its actions. Duffy emphasizes that even as highly capable models like GPT-6 Astra perform tasks with less chain-of-thought prompting, structured game harnesses remain essential to force verifiable behaviors, such as writing code instead of calculating answers mentally. Good Start Labs commercializes this by selling game trajectories and complete end-to-end learning environments to frontier labs looking for high-quality reinforcement learning data.
This is our own summary of reporting by Latent Space



