
Imagine managing your greenhouse business with an AI that not only understands your plants and customers but also makes crucial management decisions—decisions that can save thousands of euros or cost you dearly. What if you could test different AI personalities and see which one performs best under real-world pressure? That’s exactly what the live experiment from Firmulate offers—a rare glimpse into how frontier AI models handle high-stakes business crises in a realistic setting.
The Live AI Business Wargame: Testing AI Management Personalities
At the heart of this experiment is a small software company, running every weekday with real money mechanics, real crises, and real temptations. The company operates with 13 synthetic employees and a complex set of rules that guide its decisions, all staged in a watchable environment at firmulate.com/live. Each day, different AI models take on the role of management, navigating challenges like customer crises, ethical dilemmas, and strategic negotiations.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models in the Arena
Four frontier AI models competed in this live scenario, each with distinct traits. They ranged in scores from 95 to 73 out of a possible 100 in a leaderboard, reflecting their overall performance in a management context. The top scorer, GPT-5.6-sol, demonstrated exceptional ability by uncovering a critical piece of information buried in the company’s files and successfully closing a €55,000 deal. Meanwhile, the other models—Kimi K3, Sonnet 5, and Fable 5—also closed deals but with varying degrees of discipline and thoroughness.
Key Findings: Honesty and Diligence Matter
Despite their differences, all models managed to identify and refuse manipulation attempts, such as escalating fake CEO messages or a reporter trick asking for a simple yes/no confirmation. Kimi K3, for example, responded with a clear reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” But the real decisive factor was whether the AI read the company’s documentation thoroughly before making decisions.
The models that read deeper and more carefully—like GPT-5.6-sol—secured the full deal, leading to a revenue increase of +€4,583 MRR (monthly recurring revenue). Conversely, the model with the deepest analysis, Opus 4.8, despite its thoroughness, ultimately left the deal on the table, illustrating how even the best AI can slip in discipline, especially if not optimized for certain decision parameters.
The Human-Like Personalities of AI
This experiment reveals that AI models don’t just process data—they exhibit distinct management personalities. Some are thorough and cautious, refusing manipulative tactics and uncovering hidden information. Others are more prone to slip into shortcuts or leave opportunities unexplored, especially under pressure. These traits are measurable and directly impact business outcomes, making AI personality profiling an essential part of deploying automation at scale.
The Greenhouse Analogy: Why This Matters for Outdoor Business
For outdoor and greenhouse entrepreneurs, the lesson is clear: when considering AI for customer management, supply chain, or strategic planning, the key isn’t just how well it chats or how clever its responses are—it’s whether it can complete its tasks honestly and thoroughly. Just as in your greenhouse business, where a misstep can cost you plants or profits, AI must demonstrate reliability under pressure and a commitment to data integrity. The real-world stakes are high, and the experiment proves that only the most disciplined AI models succeed in delivering tangible results.
How to Test Your AI Workforce
Interested in evaluating your own AI tools? Firmulate offers a unique pilot program where enterprises can run the same management wargame against a read-only export of their business data—nothing ever writes back, ensuring safety while revealing true capabilities. Whether you’re managing a greenhouse supply chain or a tech startup, this approach helps ensure your AI will finish what it starts and uphold your standards before you fully adopt it. Learn more at firmulate.com/pilot.html.

In a real-world live experiment, frontier AI models proved capable of handling crises, refusing manipulation, and closing profitable deals—traits vital for trustworthy business automation. The key takeaway for outdoor business owners: evaluate your AI not just on its chat skills but on its ability to finish tasks honestly and thoroughly under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html