firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

AI automation meets a balance sheet

Most AI tools are demonstrated in controlled settings: a polished prompt, a clean output and no lasting consequences. Firmulate offers a harsher picture. Its software company has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and displays a public cash countdown. The business is not presented as a hypothetical. Its work continues, every workday is versioned, and the struggle can be watched live.

For readers interested in AI tools and automation, that makes Firmulate unusually revealing. It asks what happens after an agent produces a plausible answer. Will it inspect the relevant files, withstand pressure, complete the commercial task and preserve trust? The company’s precarious finances turn those questions into an unfolding business story rather than another chatbot showcase.

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company learning in public

Firmulate describes its product as an AI company emulator, but the live company is also a running portrait of automated management under pressure. Its synthetic workforce has accumulated more than 680 self-learned playbook rules. Each workday creates another auditable record of what the company noticed, decided and actually completed.

That record matters because the experiment is built around real money mechanics. With burn of €105k a month and only €2.3k MRR, the public countdown gives routine work an unmistakable context: the company must improve its position. Visitors can follow the operation itself and read what its synthetic employees say, making progress and failure visible as they happen.

Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results

Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, repeated fairly

The Crucible League put frontier models through the same small software company during its worst week. They faced the same customers, crises and temptations, while every decision was versioned and auditable. The final July 2026 standings placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.

The headline result was not that the models failed to understand the situation. All of them spotted every crisis, and all refused every manipulation attempt. The decisive gap appeared between analysis and execution: only two signed the €55,000 deal that their own work had earned. Firmulate summarized the problem in six words: “Same diagnosis, same pitch — no signature.”

The deal also exposed the practical value of reading company material carefully. A decisive weakness in the competitor’s position was not placed in the customer event. It sat two document references deep in the company’s own files. The models that found it won the deal at full price, adding €4,583 MRR. In this test, retrieving buried business context was not clerical diligence; it changed the commercial outcome.

Trust survived the pressure test

The models also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is significant because the league treats trust as a hard boundary. A single breach caps the total, under the principle that “no amount of good work outweighs a breach of trust.” The models protected that boundary even while other parts of their performance varied sharply.

Thoroughness was not enough

Opus 4.8 illustrates why evaluating automation through eloquence or effort alone can mislead. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. Yet it finished last. It failed to close the deal and repeatedly attempted to write into a locked department instead of escalating the blockage.

The same discipline problem appeared more weakly in the other four participants. Firmulate’s portrait is therefore not a simple contest between capable and incapable systems. It shows models recognizing danger, behaving honestly and still leaving essential work unfinished. Kimi K3’s result also carries an important fairness note: it ran with the API default because it had no effort parameter, while the others ran at xhigh.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Workflow Automation with Microsoft Power Automate: Use business process automation to achieve digital transformation with minimal code

Workflow Automation with Microsoft Power Automate: Use business process automation to achieve digital transformation with minimal code

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Automation should be judged by what reaches the finish line

Firmulate’s live company turns build-in-public into a continuous management test. Its synthetic employees do not merely generate text; they operate inside a business with customers, documents, rules, revenue pressure and a visible countdown. That makes every missed close, careful refusal and newly learned rule part of the company’s public history.

The lesson for organizations adopting AI tools is direct. Spotting a crisis is not the same as resolving it. Producing a strong pitch is not the same as securing the signature. Reading the obvious event is not the same as finding the fact buried in company files. And a large body of thoughtful work does not compensate for weak follow-through.

Firmulate makes those gaps watchable rather than theoretical. As the company continues to burn €105k a month against €2.3k MRR, its survival story supplies fresh evidence about whether automated workers can turn sound judgment into completed, trustworthy business outcomes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Seedance 2.5

Seedance 2.5 has been officially released, introducing significant updates to the platform. Details on features, impact, and future plans are outlined.

Jack Dorsey Criticizes Tether for Limited Donation to Opensats

Ongoing criticism from Jack Dorsey raises questions about Tether’s true intentions and the future of decentralized blockchain innovation.

As AI Divides Chile’s Leaders, the Nation Becomes a Mirror of Global Indecision.

Just as Chile’s leaders debate AI, the world faces a crossroads that could redefine innovation and ethics—discover how this global dilemma unfolds.

Fasset Obtains Malaysia License to Launch First Islamic Digital Bank

L isten to how Fasset’s Malaysia license for the first Islamic digital bank could reshape ethical banking and your financial future.