
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Score of 26 for Doing Nothing — On Purpose
If you’ve ever wondered whether AI benchmark scores actually mean anything, here’s a refreshing oddity: in Firmulate’s benchmark league, a manager that does literally nothing still walks away with 26 points. Not zero. Twenty-six.
That number is deliberate. It’s the floor of a scoring philosophy that says something uncomfortable about most AI evaluations: if a do-nothing baseline earns 26, then every point above it represents work that was actually finished, files that were actually read, and trust that was actually kept. It’s a benchmark built for people who plan to put AI agents near their CRM, support queue, or forecast — and want to know what happens on the worst week, not the best demo.
AI decision-making support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crises, Only the Model Changes
Firmulate ran four frontier AI models through an identical scenario: each was put in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing about a model’s performance is anecdotal.
The Crucible League’s final July 2026 standings tell the story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.)
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Counts — and Why 26 Isn’t Zero
The do-nothing baseline scores 26 because the benchmark rewards partial progress. A manager that identifies a crisis but doesn’t resolve it still did something real. Spotting the problem, containing the damage, making the right diagnosis — those are units of genuine work, and the floor of 26 reflects what competent-but-incomplete management looks like.
But there’s a ceiling rule too, and it’s the more interesting one: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model could ace every operational decision and still see its score capped by one moment of dishonesty. That’s a standard most human managers aren’t held to.
As an affiliate, we earn on qualifying purchases.
The Finding That Chat Demos Can’t Show
Here’s what separates the top of the table from the bottom. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The others simply left the close on the table.
The decisive detail was buried. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant in the field: over 80 learned rules and the deepest analyses of any model. And it finished last. The close went unsigned, and discipline slipped — it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
As an affiliate, we earn on qualifying purchases.
Pressure-Testing Honesty
The experiment didn’t just test competence. It tested nerve. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want documented before an agent touches real systems, not after.
You Can Watch It Live
Firmulate isn’t a static leaderboard. It runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and you can watch it at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via the pilot program at firmulate.com/pilot.html.

The Takeaway
An honest benchmark has strange-looking numbers by design: a floor at 26 that makes partial progress visible, a trust cap that makes honesty non-negotiable, and a healthy suspicion of tidy round 100s. When you’re evaluating AI tools for your business, that’s the standard worth demanding — because the gap between “writes beautifully” and “finishes the job” is exactly where your revenue lives. The €55,000 deal nobody signed proves it.
See the full results and plain-language findings at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
