
If you’re choosing an AI agent to handle customer support, sales or operations, a polished demo leaves a big question unanswered: will it finish the job when the week gets difficult? In Firmulate’s company-management experiment, Moonshot’s Kimi K3 finished just behind the leader—and ahead of three of four Western frontier models.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A crowded field, a narrow win
In the final July 2026 Crucible League, gpt-5.6-sol placed first with 95 points. Kimi K3 came second at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”
Firmulate put each frontier model in charge of the same small software company during its worst week, with identical customers, crises and temptations. Its claim is that decisions are versioned and auditable. The test asks how models manage a business, rather than how well they respond in a chat demo.
As an affiliate, we earn on qualifying purchases.
Spotting the problem wasn’t enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing an opportunity and acting on it is the experiment’s sharpest finding: “Same diagnosis, same pitch — no signature.”
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that security-related clue, closed the deal, saved the churning customer and resisted all three baits. Firmulate reports just one deviation for K3, the cleanest discipline in the field.
As an affiliate, we earn on qualifying purchases.
Restraint under pressure, follow-through under scrutiny
The manipulation test included fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a concrete example of the kind of boundary an agent may need to hold when a request appears to come from someone with authority.
Opus 4.8 offers a counterpoint to the idea that more activity automatically means better management. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says a weaker version of that same weakness appeared in all four.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A live test, with a fairness caveat
The company behind the benchmark is a live experiment, not a fictional case study. Firmulate describes 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the company can be watched at Firmulate.
There is a relevant caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the league result when interpreting the rankings.
Firmulate also says 242 real, unedited management decisions power a “guess the model” quiz. For enterprises, it offers a pilot using a read-only export of a company’s business; the stated arrangement does not write back to real systems. The benchmark page lays out the results.

AI security and compliance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the work, not the demo
K3’s second-place finish is close enough to the leader to make the result interesting, and strong enough to put three Western frontier models behind it. But the wider lesson is about evaluating agents on the work you expect them to do: reading relevant records, closing a deal, protecting customers and resisting pressure. Choosing a model without testing it on your own tasks is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
