firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

If you’re choosing an AI agent to handle customer support, sales or operations, a polished demo leaves a big question unanswered: will it finish the job when the week gets difficult? In Firmulate’s company-management experiment, Moonshot’s Kimi K3 finished just behind the leader—and ahead of three of four Western frontier models.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A crowded field, a narrow win

In the final July 2026 Crucible League, gpt-5.6-sol placed first with 95 points. Kimi K3 came second at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”

Firmulate put each frontier model in charge of the same small software company during its worst week, with identical customers, crises and temptations. Its claim is that decisions are versioned and auditable. The test asks how models manage a business, rather than how well they respond in a chat demo.

Amazon

AI customer support chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Spotting the problem wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing an opportunity and acting on it is the experiment’s sharpest finding: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that security-related clue, closed the deal, saved the churning customer and resisted all three baits. Firmulate reports just one deviation for K3, the cleanest discipline in the field.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Restraint under pressure, follow-through under scrutiny

The manipulation test included fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a concrete example of the kind of boundary an agent may need to hold when a request appears to come from someone with authority.

Opus 4.8 offers a counterpoint to the idea that more activity automatically means better management. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says a weaker version of that same weakness appeared in all four.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A live test, with a fairness caveat

The company behind the benchmark is a live experiment, not a fictional case study. Firmulate describes 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned, and the company can be watched at Firmulate.

There is a relevant caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the league result when interpreting the rankings.

Firmulate also says 242 real, unedited management decisions power a “guess the model” quiz. For enterprises, it offers a pilot using a read-only export of a company’s business; the stated arrangement does not write back to real systems. The benchmark page lays out the results.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI security and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not the demo

K3’s second-place finish is close enough to the leader to make the result interesting, and strong enough to put three Western frontier models behind it. But the wider lesson is about evaluating agents on the work you expect them to do: reading relevant records, closing a deal, protecting customers and resisting pressure. Choosing a model without testing it on your own tasks is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is The Defender’s Window The Key To AI’s Safety And Ethics?

OpenAI warns organizations have limited time to deploy AI-based cybersecurity defenses before attackers gain similar capabilities, outlining a four-part strategy.

Building The Materials Foundation For AI

Emerging focus on advanced materials is shaping the future of AI hardware, with increasing research interest and ongoing development efforts.

We Opened A Taco Shop Because Meta Is Building An AI Data Center Nearby. Now 40% Of Our Business Is Tied To It.

A local taco shop reports that 40% of its sales now come from customers connected to Meta’s nearby AI data center project, following its opening near the site.

How Affordable AI Is Redefining The Rules Of The Open-Weight Competition

Alibaba’s release of the low-cost, open-licensed Qwen3.8-Flash-Next is reshaping AI distribution and competition, emphasizing efficiency over raw power.