
A greenhouse can look calm right up until a pump fails, a supplier misses delivery and a customer wants an answer. If AI agents are going to help run a business, the useful question is how they handle that kind of pressure—not how polished they sound in a demo.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts AI models through a live, watchable company experiment. Its enterprise pilot takes the idea to a business’s own data: test crisis decisions and playbooks before agents touch real operations.
A company’s worst week, repeated
In the final Crucible League, published in July 2026, each frontier model ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was captured in the experiment’s summary: “Same diagnosis, same pitch — no signature.”
The detail hidden in the files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that sound advice depends on finding and using relevant information, then following through.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The experiment measured more than whether a model could identify a threat; it showed whether it held a boundary when the request was framed to feel routine.
Thorough work still needs a close
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The deal was left unsigned, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The leaderboard is a result from this experiment, not a universal prediction of how those models will perform in every business.
From watching to a company-specific pilot
The live company makes the stakes visible. It has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. Readers can watch the experiment at Firmulate. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
For an enterprise, the next step is a wargame using a read-only export of its own business. Teams can test crisis scenarios against their customers, pipeline, rules and playbooks, then review a board report with model rankings and weak points. The pilot is designed so nothing writes back to real systems. That turns a public experiment into a chance to examine how an AI workforce might respond to the specific pressures a company faces.

Try the pilot
To explore a wargame against your company’s read-only data, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
