Build A Tough Trial Week For AI Agents Before Business Launch
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Build A Tough Trial Week For AI Agents Before Business Launch on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate’s final Crucible League, completed in July 2026, ran five frontier AI models through a simulated company’s worst week. All models detected every crisis and refused every manipulation attempt, but only two closed a €55,000 deal their own analysis had earned. Firmulate now offers enterprise pilots using read-only exports of a company’s own data.

Firmulate has completed the final round of its Crucible League, a live experiment that ran five frontier AI models through the worst week of the same small software company, and is now offering enterprise pilots that apply the same test to a read-only export of a company’s own data. The July 2026 experiment, run by Thorsten Meyer AI and watchable at firmulate.com, produced a clear split: every model spotted every crisis and refused every manipulation attempt, but only two of five signed the €55,000 deal their own analysis had justified.

The final standings placed gpt-5.6-sol first at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted toward scores, but a single breach of trust capped the total under a rule the experiment summarized as “no amount of good work outweighs a breach of trust.” Every decision was versioned and auditable.

The decisive gap between models was not crisis detection but follow-through. According to the experiment’s findings, models that read a key piece of evidence buried two document references deep in the company’s own files went on to close the deal at full price, worth +€4,583 in monthly recurring revenue. The experiment described the failure mode as: “Same diagnosis, same pitch — no signature.” In other words, an agent can recognize the situation and make a persuasive case yet still fail to act on information already available inside the business.

Trust was tested separately. Fake CEO messages escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

At a glance
reportWhen: league completed July 2026; enterprise…
The developmentFirmulate completed its final Crucible League in July 2026 and is now offering enterprise pilots that run the same wargame format against read-only exports of real company data.
Build A Tough Trial Week For AI Agents Before Business Launch
Crucible League · Final Round · July 2026

Build a Tough Trial Week for AI Agents Before Business Launch

Firmulate ran five frontier AI models through a simulated company’s worst week. Every crisis was detected and every manipulation refused — but only two of five closed the €55,000 deal their own analysis had earned. The gap wasn’t diagnosis. It was follow-through.

5 / 5
Models detected every crisis & refused every manipulation
2 / 5
Closed the justified €55,000 deal at full price
+€4,583
Monthly recurring revenue tied to the buried-evidence deal
13
Synthetic employees
€105k
Monthly burn
680+
Self-learned playbook rules
242
Real decisions in the public quiz
€2.3k
Baseline MRR
Final Standings

One Winner, One Clear Split

gpt-5.6-sol took first place at 95 points. Partial progress counted toward scores, but a single breach of trust capped the total. Every decision was versioned and auditable.

RankModelScoreEffort SettingCrises DetectedRefused ManipulationClosed the Deal
01gpt-5.6-sol95xhigh✓ All✓ Yes✓ Yes
02Kimi K393API default✓ All✓ Yes~ Partial
03Sonnet 588xhigh✓ All✓ Yes~ Partial
04Fable 577xhigh✓ All✓ Yes✗ No
05Opus 4.873xhigh✓ All✓ Yes✗ No
—Do-nothing baseline26n/a✗ None✗ n/a✗ No

Fairness caveat: Kimi K3 ran without an effort parameter (API default), while the other four models ran at xhigh. ThorstenMeyerAI.com notes the standings are a record of this experiment, with that configuration difference as context for interpretation.

Score Visualization

The Field at a Glance

Scores from the July 2026 final round. The do-nothing baseline shows how much the worst week punishes inaction on its own.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
Why Diagnosis Without Closure Matters

The Hardest Failure to Catch in a Demo

The decisive gap was not crisis detection but follow-through. Models that read a key piece of evidence buried two document references deep went on to close the deal at full price. A weaker version of that weakness appeared in all five models.

Evidence

Find Buried Information

The winning models dug two references deep into the company’s own files to surface the justification for the €55,000 opportunity — worth +€4,583 in MRR.

Execution

Close What You Earned

Three models made the same diagnosis and the same persuasive pitch, yet never acted. The experiment’s verdict: “Same diagnosis, same pitch — no signature.”

Boundaries

Respect Locked Doors

Opus 4.8 — the most thorough participant with 80 learned rules — attempted to write into a locked department instead of escalating, and left the deal unclosed.

On the Record

Three Lines That Defined the Experiment

“No amount of good work outweighs a breach of trust.”

Experiment rules — ThorstenMeyerAI.com

“Same diagnosis, same pitch — no signature.”

Findings — ThorstenMeyerAI.com

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 — on-record reasoning
From Synthetic Company to Your Data

The Wargame Behind the Standings

The Trust Gauntlet

Fake CEO messages escalated over three stages, followed by a reporter’s request framed as “just one yes/no, on background.” All five models refused every attempt.

3 stagesEscalating fake-CEO pressure
1 trapReporter’s “on background” yes/no request
5 / 5Models refused every manipulation

The Synthetic Company

A 13-employee software firm with real money mechanics, a public cash countdown, and versioned workdays. Readers can follow it live at firmulate.com.

€105k / moBurn against €2.3k in recurring revenue
680+ rulesSelf-learned playbook, grown in-simulation
242 callsUnedited decisions in the public quiz
Running Your Own Wargame

How an Enterprise Pilot Works

The pilot extends the same format to a company’s own operations. Nothing writes back to real systems. Full benchmark results: firmulate.com/benchmarks.html.

1

Read-Only Export

🔐 A one-way export of your company’s data — no write access to live systems.

2

Crisis Scenarios

⚡ The wargame runs your worst week: crises, traps, and a buried opportunity.

3

Model Rankings

📊 Agents are scored under the no-breach-of-trust cap, every decision versioned.

4

Board Report

📋 Rankings plus identified weak points in your own playbooks. Arrange via contact@firmulate.com.

Key Questions

What Readers Ask Most

What is the Crucible League?

A live experiment run by Thorsten Meyer AI in which frontier AI models each ran the same simulated small software company through its worst week. The final round was completed in July 2026, with every decision versioned and auditable.

Which model won, and by how much?

gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26.

Did any of the AI models fall for the fake CEO messages?

No. All five models refused the escalating fake CEO messages and a reporter’s background request. Kimi K3 explicitly classified the request as a suspected approval-bypass or impersonation.

Is the model comparison entirely fair?

Not entirely. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results also come from a single synthetic company in a single simulated week — generalization to other businesses is not established by this experiment.

How can a company test its own business this way?

Through a Firmulate enterprise pilot, which runs a wargame against a read-only export of the company’s own data. No data writes back to real systems, and the output is a board report with model rankings and identified playbook weaknesses. The live simulation and decision quiz remain at firmulate.com/live.

Why Diagnosis Without Closure Matters

The results point to a practical problem for businesses adopting AI agents: the hardest failure to catch in a demo is the one where the agent understands the situation but does not act on it. Firmulate’s league suggests that spotting a crisis and refusing a scam are not the whole job. Agents also need to find relevant evidence, close a justified opportunity, and respect boundaries when their first route is blocked.

That last behavior also separated the field. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It attempted to write into a locked department instead of escalating, and left the deal unclosed. A weaker version of that same weakness appeared in all four other models, according to the experiment’s findings. For companies weighing automation, the lesson is that thoroughness and correctness in analysis do not automatically translate into disciplined execution.

From Synthetic Company to Your Data

Firmulate’s live simulation is built around a synthetic software company with 13 employees and real money mechanics: a burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and versioned workdays. Readers can follow the simulation at firmulate.com, including a quiz built from 242 real, unedited management decisions that invites readers to guess which model made each choice.

The enterprise pilot extends the same format to a company’s own operations. A pilot uses a read-only export of the company’s data to test crisis scenarios and produce a board report with model rankings and identified weak points in the company’s own playbooks. Nothing writes back to real systems, according to ThorstenMeyerAI.com. Full benchmark results are published at firmulate.com/benchmarks.html.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s experiment rules, per ThorstenMeyerAI.com

Caveats in the Standings

The comparison carries a stated fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other four models ran at the xhigh effort setting. ThorstenMeyerAI.com notes that the standings are a record of this experiment and that the configuration difference is part of the context for interpreting them.

Beyond that, the results come from a single synthetic company in a single simulated week. How far the findings generalize to other business types, larger organizations, or live operational environments is not established by this experiment. The enterprise pilot is positioned as the mechanism for testing a specific company’s scenarios rather than a claim about universal model performance.

Running Your Own Wargame

Companies interested in testing the format against their own operations can arrange a pilot through Firmulate’s pilot page or by contacting contact@firmulate.com. Each pilot uses a read-only data export, runs crisis scenarios against it, and delivers a board report with model rankings and playbook weak points. The live simulation and the decision quiz remain available at firmulate.com/live for readers who want to observe the mechanics before committing to a pilot.

Source: ThorstenMeyerAI.com

Key Questions

What is the Crucible League?

A live experiment run by Thorsten Meyer AI in which frontier AI models each ran the same simulated small software company through its worst week. The final round was completed in July 2026, with every decision versioned and auditable.

Which model won, and by how much?

gpt-5.6-sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26.

Did any of the AI models fall for the fake CEO messages?

No. All five models refused the escalating fake CEO messages and a reporter’s background request. Kimi K3 explicitly classified the request as a suspected approval-bypass or impersonation.

Is the model comparison entirely fair?

Not entirely. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. ThorstenMeyerAI.com states the standings are a record of this experiment with that difference as context.

How can a company test its own business this way?

Through a Firmulate enterprise pilot, which runs a wargame against a read-only export of the company’s own data. No data writes back to real systems, and the output is a board report with model rankings and identified playbook weaknesses.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why SpaceXAI’s Recent Cursor Acquisition Is A Game-Changer For AI Development

SpaceXAI has completed its acquisition of Cursor after releasing Grok Bot and Grok 4.6, potentially expanding its AI tools for software development.

Exploring The AI Tower: Twelve Rooms Of Safe And Practical AI Application

A detailed look at the new AI Tower, featuring twelve practical rooms demonstrating safe, effective AI uses without tracking or sign-up. Key insights and future steps.

The 13 Most Effective AI Tools For Student Productivity In 2026

Discover the 13 most effective AI tools for student productivity in 2026, including features, benefits, and how they enhance learning efficiency.

Proaction’s Codex Story: Higher Sales And 75+ Hours Saved

OpenAI’s customer story says Brazilian cosmetics company Proaction raised sales 60% and saved 75+ hours with Codex; neither figure is independently verified.