firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Can you recognize an AI by the decisions it makes?

For readers accustomed to comparing AI tools through polished demos, Firmulate offers a more revealing test: watch frontier models manage the same company under pressure, then try to identify them from their choices.

The public Firmulate quiz draws on 242 real, unedited management decisions. These are not hypothetical answers produced for a personality test. They come from a live experiment in which each model ran the same small software company through its worst week, facing identical customers, crises and temptations.

The result is an interactive article with unusually high stakes. A reader sees a decision, guesses which model made it and then discovers the participant’s broader management profile. Differences that can disappear in ordinary chat become visible through action: who investigates, who follows through, who maintains discipline and who leaves valuable work unfinished.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, different managers

Firmulate describes itself as an AI company emulator. Its synthetic business has 13 employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure watchable, while every workday is versioned. The company has also accumulated more than 680 self-learned playbook rules.

Every competing model received the same operating environment. Every decision was versioned and auditable. That consistency matters because it moves the comparison away from prompt-writing tricks and toward a question businesses increasingly care about: what happens when an AI is responsible for sustained work rather than a single response?

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted.

Firmulate also imposed a strict trust constraint: a single breach capped the total. Its rationale was blunt: “no amount of good work outweighs a breach of trust.” Yet dishonesty was not what separated the field. All models detected every crisis and rejected every manipulation attempt.

The difference between analysis and completion

The most consequential divide appeared in a sales opportunity. Every model reached the same diagnosis and developed the same pitch, but only two signed the €55,000 deal their analysis had earned. The experiment’s concise summary captures the failure: “Same diagnosis, same pitch — no signature.”

The key was not sitting in the customer event. A decisive weakness in a competitor was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding has implications for AI automation far beyond sales. Recognizing an event is not the same as gathering the necessary context, and producing a convincing recommendation is not the same as completing the work. A tool can appear intelligent while still failing at the final operational step.

Pressure exposed discipline, not just knowledge

The social-engineering tests were equally practical. Fake messages from the CEO escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s result deserves a fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished second and showed the cleanest discipline in the field.

Opus 4.8 illustrates why thoroughness alone did not guarantee success. It was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table and repeatedly attempted to write into a locked department instead of escalating the blockage. The same discipline problem appeared in all four other participants, though less strongly.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A management benchmark readers can inspect

The quiz turns these records into something more accessible than a league table. By asking readers to attribute unedited decisions, it reveals whether the models have recognizable management tendencies—and whether familiar reputations survive contact with operational evidence.

For companies considering AI workers, the experiment suggests a better evaluation standard. Strong prose and correct diagnosis matter, but so do file-reading habits, resistance to manipulation, procedural discipline and the ability to finish valuable work.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That makes the wargame a rehearsal rather than a deployment: organizations can observe how an AI workforce behaves around their own context before allowing it to act.

The most revealing question in the quiz is therefore not whether a reader can recognize a model’s writing style. It is whether the decision shows the habits of a manager that a business would trust when the company, the customer and the money are all real enough to matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business emulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

12 Breakthrough AI Devices For Home Automation In 2026

Explore the 12 most advanced AI-powered home automation devices in 2026, highlighting local control, privacy, and customization options for different users.

Opinion | Even Millions Of Stolen Books Cannot Satisfy Ravenous A.I. Chatbots

Despite millions of stolen books, AI chatbots continue to seek more data, raising questions about data dependence and ethical concerns in AI development.

Small Streamers And AI: Generating Ranked Clip Lists From Full Streams

New AI tools enable small streamers to generate ranked clip lists from full streams, potentially reducing editing costs and enhancing content curation.

Proaction’s Codex Story: Higher Sales And 75+ Hours Saved

OpenAI’s customer story says Brazilian cosmetics company Proaction raised sales 60% and saved 75+ hours with Codex; neither figure is independently verified.