firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Can you recognize an AI by the decisions it makes?

For readers accustomed to comparing AI tools through polished demos, Firmulate offers a more revealing test: watch frontier models manage the same company under pressure, then try to identify them from their choices.

The public Firmulate quiz draws on 242 real, unedited management decisions. These are not hypothetical answers produced for a personality test. They come from a live experiment in which each model ran the same small software company through its worst week, facing identical customers, crises and temptations.

The result is an interactive article with unusually high stakes. A reader sees a decision, guesses which model made it and then discovers the participant’s broader management profile. Differences that can disappear in ordinary chat become visible through action: who investigates, who follows through, who maintains discipline and who leaves valuable work unfinished.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, different managers

Firmulate describes itself as an AI company emulator. Its synthetic business has 13 employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure watchable, while every workday is versioned. The company has also accumulated more than 680 self-learned playbook rules.

Every competing model received the same operating environment. Every decision was versioned and auditable. That consistency matters because it moves the comparison away from prompt-writing tricks and toward a question businesses increasingly care about: what happens when an AI is responsible for sustained work rather than a single response?

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted.

Firmulate also imposed a strict trust constraint: a single breach capped the total. Its rationale was blunt: “no amount of good work outweighs a breach of trust.” Yet dishonesty was not what separated the field. All models detected every crisis and rejected every manipulation attempt.

The difference between analysis and completion

The most consequential divide appeared in a sales opportunity. Every model reached the same diagnosis and developed the same pitch, but only two signed the €55,000 deal their analysis had earned. The experiment’s concise summary captures the failure: “Same diagnosis, same pitch — no signature.”

The key was not sitting in the customer event. A decisive weakness in a competitor was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding has implications for AI automation far beyond sales. Recognizing an event is not the same as gathering the necessary context, and producing a convincing recommendation is not the same as completing the work. A tool can appear intelligent while still failing at the final operational step.

Pressure exposed discipline, not just knowledge

The social-engineering tests were equally practical. Fake messages from the CEO escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result deserves a fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished second and showed the cleanest discipline in the field.

Opus 4.8 illustrates why thoroughness alone did not guarantee success. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table and repeatedly attempted to write into a locked department instead of escalating the blockage. The same discipline problem appeared in all four other participants, though less strongly.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A management benchmark readers can inspect

The quiz turns these records into something more accessible than a league table. By asking readers to attribute unedited decisions, it reveals whether the models have recognizable management tendencies—and whether familiar reputations survive contact with operational evidence.

For companies considering AI workers, the experiment suggests a better evaluation standard. Strong prose and correct diagnosis matter, but so do file-reading habits, resistance to manipulation, procedural discipline and the ability to finish valuable work.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That makes the wargame a rehearsal rather than a deployment: organizations can observe how an AI workforce behaves around their own context before allowing it to act.

The most revealing question in the quiz is therefore not whether a reader can recognize a model’s writing style. It is whether the decision shows the habits of a manager that a business would trust when the company, the customer and the money are all real enough to matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business emulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Claude Users, Here’s Why Anthropic’s Auto Mode Is A Game-Changer

Anthropic announces auto mode will become the default setting for Claude starting August 14, but details on operation and affected products remain unclear.

AI Tools & Automation: The Complete Hub for Building a Smarter Digital Life

AIThis post was created with the assistance of artificial intelligence (AI).Artificial intelligence…

DeepSeek V4 Pro’s AI Leap: Is The 5% Improvement Justified At 4,500% More Cost?

DeepSeek claims its V4 Pro model is upgraded, with a purported 5% performance gain over Claude at 4,500% higher cost. Details remain unverified.

Qwen3.8-2.4T

Qwen3.8-2.4T is a new AI model announced, featuring 3.8 billion parameters and 2.4 trillion tokens, marking a significant step in AI development.