
Imagine testing a new gardening tool by trying to grow a plant with it — and discovering it doesn’t do anything at all. Now, what if that tool still earned some points for simply being there? That’s the essence of a recent AI benchmark experiment that evaluates how well artificial intelligence models handle real-world decision-making — even when they do nothing.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark: More Than Just Chat Quality
In a groundbreaking experiment, four advanced AI models were tasked with managing a simulated small software company during its most challenging week. This wasn’t just about generating text or answering questions — the models had to navigate crises, read critical files, and make decisions that a human manager would face. The goal was to assess their management quality, not their conversational prowess.
Each AI was placed in the same scenario, dealing with identical customer issues, internal conflicts, and sales opportunities. Every decision was tracked, versioned, and auditable — making the process transparent and comparable. The experiment provides a glimpse into the operational integrity and reliability of AI models in complex, real-world tasks.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline: Why Do-Nothing Scoring Starts at 26
One of the most striking findings is that even a completely inactive AI — a ‘do-nothing’ baseline — scores 26 points out of a possible 100. This might seem counterintuitive, but it reflects the benchmark’s design: partial progress is meaningful, and the scoring system recognizes the value of even minimal engagement.
Moreover, the system enforces a strict rule: if an AI breaches trust — for example, by attempting manipulation or deception — its total score is capped, regardless of other successes. This ensures that honesty and integrity are prioritized over mere task completion. Essentially, the full score isn’t just about what the AI does, but what it refuses to do.
The Key to Winning: Reading the Critical Files
While all models identified crises and refused manipulation attempts, only two managed to secure the deal worth €55,000 — the equivalent of a significant monthly recurring revenue. The secret lay in reading two document references deep into the company’s files. Those models that examined the hidden information discovered crucial insights, leading to the successful deal closure at full price, worth over €4,583 monthly recurring revenue.
Trust and Integrity: The Ultimate Score Cap
Trustworthiness is a core component of the benchmark. When models faced social engineering attempts — fake CEO messages escalating over stages, or a reporter asking for a simple ‘yes/no’ on background — all five models refused. Kimi K3’s on-record reasoning was clear: ‘Treat the request as a suspected approval-bypass / possible impersonation.’
This strict stance underscores that honesty and resistance to manipulation are valued more than simply completing tasks. A breach of trust, even a small one, caps the total score, emphasizing that integrity cannot be sacrificed for short-term gains.
Real-world AI in Action: The Live Company Environment
The experiment isn’t just theoretical. It involves a live setup mimicking a small company with 13 synthetic employees and real money mechanics — burning €105,000 each month against €2,300 in monthly recurring revenue. The environment is transparent, versioned daily, and accessible for observation at firmulate.com/live. This setup shows how AI models perform under pressure, with real financial stakes and operational complexity.
The Performance of Leading Models: Who Comes Out on Top?
- GPT-5.6-sol scored 95, found the buried fact, and closed the deal — delivering complete performance.
- Kimi K3 scored 93, also closed the deal, and demonstrated the cleanest discipline, running without an effort parameter.
- Sonnet 5 scored 88, closed the deal but with slightly more process slips.
- Opus 4.8 scored 77, managed to close but with notable process weaknesses, such as leaving the close on the table or escalating instead of resolving internally.
Why This Matters for Business
In real-world settings — whether managing customer relationships, forecasting, or support — the question isn’t just about AI’s ability to generate convincing text. It’s whether AI can finish what it starts, read critical information, and stay honest under pressure. These qualities are essential for trust and operational effectiveness.
Takeaway: A Benchmark That Values Trust and Results
This experiment sets a new standard for AI evaluation. It recognizes that partial success isn’t enough — reading hidden critical data, refusing manipulation, and maintaining trust are foundational. And the fact that the baseline score is 26 points, even when the AI does nothing, underscores that the framework values honesty as much as performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
