firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Score of 26 for Doing Nothing — On Purpose

If you’ve ever wondered whether AI benchmark scores actually mean anything, here’s a refreshing oddity: in Firmulate’s benchmark league, a manager that does literally nothing still walks away with 26 points. Not zero. Twenty-six.

That number is deliberate. It’s the floor of a scoring philosophy that says something uncomfortable about most AI evaluations: if a do-nothing baseline earns 26, then every point above it represents work that was actually finished, files that were actually read, and trust that was actually kept. It’s a benchmark built for people who plan to put AI agents near their CRM, support queue, or forecast — and want to know what happens on the worst week, not the best demo.

Amazon

AI decision-making support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Only the Model Changes

Firmulate ran four frontier AI models through an identical scenario: each was put in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing about a model’s performance is anecdotal.

The Crucible League’s final July 2026 standings tell the story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.)

Amazon

AI CRM automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Partial Progress Counts — and Why 26 Isn’t Zero

The do-nothing baseline scores 26 because the benchmark rewards partial progress. A manager that identifies a crisis but doesn’t resolve it still did something real. Spotting the problem, containing the damage, making the right diagnosis — those are units of genuine work, and the floor of 26 reflects what competent-but-incomplete management looks like.

But there’s a ceiling rule too, and it’s the more interesting one: a single breach of trust caps the total grade. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model could ace every operational decision and still see its score capped by one moment of dishonesty. That’s a standard most human managers aren’t held to.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Can’t Show

Here’s what separates the top of the table from the bottom. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The others simply left the close on the table.

The decisive detail was buried. The competitor weakness that unlocked the deal sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant in the field: over 80 learned rules and the deepest analyses of any model. And it finished last. The close went unsigned, and discipline slipped — it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

Amazon

AI deal closing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure-Testing Honesty

The experiment didn’t just test competence. It tested nerve. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want documented before an agent touches real systems, not after.

You Can Watch It Live

Firmulate isn’t a static leaderboard. It runs a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules. Every workday is versioned, and you can watch it at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via the pilot program at firmulate.com/pilot.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark has strange-looking numbers by design: a floor at 26 that makes partial progress visible, a trust cap that makes honesty non-negotiable, and a healthy suspicion of tidy round 100s. When you’re evaluating AI tools for your business, that’s the standard worth demanding — because the gap between “writes beautifully” and “finishes the job” is exactly where your revenue lives. The €55,000 deal nobody signed proves it.

See the full results and plain-language findings at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: Needle2: 14MB Agentic LLM For Phones, Wearables, Smart Home And Robots

Cactus releases Needle2, a 14MB agentic language model designed for phones, wearables, smart homes, and robots, enabling advanced AI functionalities on small devices.

12 Breakthrough AI Devices For Home Automation In 2026

Explore the 12 most advanced AI-powered home automation devices in 2026, highlighting local control, privacy, and customization options for different users.

Question: How Can You Own The Memory Of Your AI Coding Agents?

Hugging Face’s ‘funes’ adds local-first indexing and retrieval for AI coding agents, enabling session continuity and provenance tracking without cloud dependence.

Unlock Efficiency With Grok: Your AI Assistant For Assigning Work

xAI announces Grok can now be assigned tasks as an AI ‘teammate,’ expanding its role beyond answering prompts, though details remain unclear.