firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A brilliant answer is not the same as a well-run business

For buyers of AI tools and automation, the familiar evaluation ritual is increasingly incomplete. Coding benchmarks can show whether a model solves a defined technical problem. Chat arenas can reveal which response people prefer. Neither necessarily tells a company what happens when an agent must triage competing crises, investigate scattered evidence, resist executive pressure and carry valuable work through to completion.

That is the measurement gap exposed by Firmulate, a live experiment that asks frontier models to operate the same small software company through its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model—and the resulting quality of management.

Amazon

AI decision-making simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models entered the same corporate pressure test

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But one breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Those results deserve attention not because they create another leaderboard, but because they measure behavior that conventional leaderboards tend to omit. The models had to notice problems, decide what deserved attention, consult the company’s own material and follow through while the situation continued to evolve.

At first glance, the field appeared uniformly capable. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

That distinction should unsettle anyone selecting an agent for a CRM, support queue, sales process or forecast. Producing the right analysis is useful, but organizations ultimately depend on completed actions. An agent can sound informed at every stage and still fail to secure the outcome that justified the work.

The winning fact was buried in ordinary company material

The decisive commercial advantage did not appear in the customer event. It sat two document references deep in the company’s own files: a competitor weakness that changed the negotiating position. Models that followed those references won the deal at full price, worth +€4,583 MRR.

This is a revealing test of workplace competence. Business evidence rarely arrives as a polished prompt containing everything an agent needs. It is distributed across records, documents and prior decisions. The valuable behavior is not merely responding fluently to the latest event; it is recognizing that the visible event may be incomplete, finding the relevant context and using it at the moment of consequence.

Security discipline held under sustained pressure

The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because pressure in a company does not always resemble a malicious prompt. It can look like urgency, authority or an apparently harmless request for confirmation. In this field, the models demonstrated that they could maintain a boundary even as the approach changed.

K3’s performance also comes with an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it belongs beside it when readers compare placements.

Thoroughness did not guarantee operational success

Opus 4.8 offers the clearest warning against equating visible effort with management quality. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

That profile captures the central problem with evaluating agents through isolated answers. An impressive analysis can coexist with weak execution. A large body of learned guidance can coexist with poor escalation. The relevant question is not simply whether the model understands the situation, but whether it converts understanding into reliable, authorized action.

A company designed to make consequences visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It is burning €105k per month against €2.3k MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. Readers can watch the experiment through the public Firmulate site and review the benchmark findings.

The project also turns 242 real, unedited management decisions into a “guess the model” quiz. That exercise highlights another evaluation challenge: without a label attached, managerial judgment may be harder to attribute than a model’s writing style suggests.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI management and decision evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Agent evaluation needs a management layer

The next useful curriculum for enterprise AI is not another collection of clean prompts. It is the churn wave, the price increase, the downround and the PR crisis: situations in which priorities collide and consequences accumulate across days.

Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That is a practical model for procurement: test an agent against the organization’s actual ambiguity before granting it operational reach.

Chat quality still matters, and coding ability still matters. But neither is a substitute for reading deeply, protecting trust, escalating correctly and finishing the job. For businesses preparing to hire AI agents, management quality is becoming its own category—and it may be the category that determines whether capability turns into value.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI testing platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI behavior assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: Huzzah – A Novel Approach To Coding With AI

Huzzah, an innovative coding editor leveraging AI, was showcased on Show HN, highlighting a novel approach to software development.

Small Business Tech Boost: AI Automation Software On Sale For Labor Day

Affordable AI automation tools for small businesses are now on sale for Labor Day, enabling cost savings and efficiency boosts without technical expertise.

How Affordable AI Is Redefining The Rules Of The Open-Weight Competition

Alibaba’s release of the low-cost, open-licensed Qwen3.8-Flash-Next is reshaping AI distribution and competition, emphasizing efficiency over raw power.

2026’S Top 15 AI Platforms For Content Development

Explore the 15 leading AI platforms for content creation in 2026, highlighting features, strengths, and what makes each tool stand out for creators and businesses.