
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A brilliant answer is not the same as a well-run business
For buyers of AI tools and automation, the familiar evaluation ritual is increasingly incomplete. Coding benchmarks can show whether a model solves a defined technical problem. Chat arenas can reveal which response people prefer. Neither necessarily tells a company what happens when an agent must triage competing crises, investigate scattered evidence, resist executive pressure and carry valuable work through to completion.
That is the measurement gap exposed by Firmulate, a live experiment that asks frontier models to operate the same small software company through its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model—and the resulting quality of management.
AI decision-making simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five models entered the same corporate pressure test
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Those results deserve attention not because they create another leaderboard, but because they measure behavior that conventional leaderboards tend to omit. The models had to notice problems, decide what deserved attention, consult the company’s own material and follow through while the situation continued to evolve.
At first glance, the field appeared uniformly capable. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
That distinction should unsettle anyone selecting an agent for a CRM, support queue, sales process or forecast. Producing the right analysis is useful, but organizations ultimately depend on completed actions. An agent can sound informed at every stage and still fail to secure the outcome that justified the work.
The winning fact was buried in ordinary company material
The decisive commercial advantage did not appear in the customer event. It sat two document references deep in the company’s own files: a competitor weakness that changed the negotiating position. Models that followed those references won the deal at full price, worth +€4,583 MRR.
This is a revealing test of workplace competence. Business evidence rarely arrives as a polished prompt containing everything an agent needs. It is distributed across records, documents and prior decisions. The valuable behavior is not merely responding fluently to the latest event; it is recognizing that the visible event may be incomplete, finding the relevant context and using it at the moment of consequence.
Security discipline held under sustained pressure
The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because pressure in a company does not always resemble a malicious prompt. It can look like urgency, authority or an apparently harmless request for confirmation. In this field, the models demonstrated that they could maintain a boundary even as the approach changed.
K3’s performance also comes with an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it belongs beside it when readers compare placements.
Thoroughness did not guarantee operational success
Opus 4.8 offers the clearest warning against equating visible effort with management quality. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
That profile captures the central problem with evaluating agents through isolated answers. An impressive analysis can coexist with weak execution. A large body of learned guidance can coexist with poor escalation. The relevant question is not simply whether the model understands the situation, but whether it converts understanding into reliable, authorized action.
A company designed to make consequences visible
The live Firmulate company has 13 synthetic employees and real money mechanics. It is burning €105k per month against €2.3k MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. Readers can watch the experiment through the public Firmulate site and review the benchmark findings.
The project also turns 242 real, unedited management decisions into a “guess the model” quiz. That exercise highlights another evaluation challenge: without a label attached, managerial judgment may be harder to attribute than a model’s writing style suggests.

AI management and decision evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Agent evaluation needs a management layer
The next useful curriculum for enterprise AI is not another collection of clean prompts. It is the churn wave, the price increase, the downround and the PR crisis: situations in which priorities collide and consequences accumulate across days.
Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That is a practical model for procurement: test an agent against the organization’s actual ambiguity before granting it operational reach.
Chat quality still matters, and coding ability still matters. But neither is a substitute for reading deeply, protecting trust, escalating correctly and finishing the job. For businesses preparing to hire AI agents, management quality is becoming its own category—and it may be the category that determines whether capability turns into value.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
