firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Fluent answers are not the same as finished work

For businesses evaluating AI tools and automation, one capability may matter more than a polished demo: will the agent inspect the available evidence before it acts?

Firmulate turned that question into a live, auditable experiment. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. All of them recognized every crisis. All rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned.

The difference was not better prose or a sharper diagnosis. The decisive information was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost automatically. As Firmulate summarized the result: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A small omission with a large commercial consequence

The experiment exposes a weakness that conventional AI demonstrations rarely reveal. A model can understand an unfolding problem, devise a credible response and still fail because it does not follow the evidence far enough. In this case, the crucial weakness of a competitor was not included in the customer event. It had to be discovered by following references through the company’s documents.

That makes “reads your files before answering” a measurable business property rather than a marketing promise. For an agent working around a CRM, support queue or forecast, the valuable behavior is not merely producing a plausible recommendation. It is locating the relevant facts, carrying them into the decision and completing the action that creates value.

The final July 2026 Crucible League shows how sharply execution separated the field:

  • gpt-5.6-sol ranked first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

A do-nothing baseline scored 26 because partial progress still counts. But Firmulate applies a hard constraint to trust: a single breach caps the total, on the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmarks page.

The strongest researcher did not deliver the strongest result

Opus 4.8 offers the most instructive caution. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. The commercial close was left on the table, while its operating discipline also slipped through write attempts into a locked department instead of escalation.

A weaker version of that discipline problem appeared in all four other participants. The lesson is uncomfortable for buyers who equate visible effort with dependable performance. Thoroughness can help an agent understand a business, but analysis is not a substitute for completing the authorized task. An agent that researches extensively and then fails at the decisive handoff may still produce no commercial result.

Trust held up under direct pressure

The models performed better when tested against manipulation. Fake CEO messages escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest operational interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That shared resistance matters because the simulated company is not a simple question-and-answer test. It has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned, and every workday is versioned.

K3’s result also carries an important qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference should remain visible when comparing outcomes, even though K3 still closed the deal.

From benchmark curiosity to procurement evidence

Firmulate’s experiment suggests that businesses should test agents inside workflows that require evidence retrieval, judgment and follow-through. A useful evaluation should reveal whether an agent opens the necessary documents, respects boundaries, escalates when blocked and finishes an approved action.

The company also turns 242 real, unedited management decisions into a quiz that asks people to guess which model made each choice. More significantly for enterprise buyers, Firmulate offers a pilot using a read-only export of the customer’s own business. Nothing writes back to real systems, allowing organizations to observe agent behavior against familiar operational material without granting production control.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading belongs on the AI buying checklist

The Firmulate result is not that frontier models missed an obvious emergency. They did not. Every model saw every crisis, and every model resisted the attempted manipulation. The meaningful divide appeared later: whether the agent followed the company’s own documentary trail and converted its correct analysis into a signed deal.

For buyers of AI automation, that distinction deserves direct testing. Ask not only whether a model can explain the situation, but whether it finds the buried fact, uses it at the right moment and completes the work. In this experiment, that behavior decided a €55,000 contract and €4,583 in monthly recurring revenue. The cost of skipping the files was not a weaker answer. It was the entire deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

auditable AI agent software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI file reading automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

DeepSeek V4 Pro’s AI Leap: Is The 5% Improvement Justified At 4,500% More Cost?

DeepSeek claims its V4 Pro model is upgraded, with a purported 5% performance gain over Claude at 4,500% higher cost. Details remain unverified.

The Ultimate Guide To Using Daybreak AI Models With AWS Infrastructure

OpenAI’s cybersecurity-focused Daybreak Blue and Red models are now available to approved AWS customers through Amazon Bedrock for security testing and vulnerability research.

Bold Claims: How Grok 4.6 Elevates AI Reasoning To New Heights

SpaceXAI announces Grok 4.6, claiming it as its new flagship AI with enhanced reasoning capabilities, but technical details and benchmarks remain undisclosed.

Is Edited’s Retail Dataset The Key To Smarter AI Workspaces?

Edited plans to embed its retail dataset into AI environments, potentially transforming retail analysis. Details on the dataset’s scale and deployment are still emerging.