🔍 Read the full analysis: Exploring Failures In Diligent AI Systems on ThorstenMeyerAI.com
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
Create a free accountAs an affiliate, we earn on qualifying purchases.
TL;DR
An ongoing live test by Firmulate demonstrates that highly diligent AI systems, despite deep analysis and strong security judgments, often fail to complete decisive business actions. The experiment underscores the gap between understanding and execution in AI automation.
Recent results from a live AI automation experiment conducted by Firmulate reveal that even the most diligent AI models can fail to complete critical business actions, despite demonstrating deep understanding and resistance to manipulation. The experiment highlights a persistent gap between problem recognition and operational execution, which has significant implications for AI deployment in real-world business processes.
In the ongoing Firmulate experiment, the AI model Opus 4.8 was the most thorough participant, producing extensive analyses and learning 80 additional playbook rules. Despite this, it finished last in a competitive evaluation, scoring only 73 points out of a possible higher total, primarily because it failed to close a key business deal. The model identified crises, resisted manipulation attempts, and developed strategies to win a major customer contract, but ultimately did not execute the final step of closing the deal.
All models involved in the experiment could recognize crises and reject manipulative requests, such as fake CEO messages, with high accuracy. However, only two models successfully signed the deal, which was supported by a critical piece of information buried two document references deep in the company’s files. This fact, when used, resulted in a deal worth €55,000 monthly recurring revenue, compared to Opus 4.8’s €4,583. The experiment underscores that the failure was not due to lack of intelligence but a breakdown in operational discipline—specifically, the inability to prioritize and act on the most decisive information.
The broader finding is that capable AI systems tend to focus on expanding their understanding and analysis but often neglect the final, impactful step of execution. Opus 4.8 learned aggressively, but its effort was spread across many rules and analyses, leading it to attempt to write into locked departments rather than escalate issues when needed. This pattern was observed across all models tested, indicating a systemic challenge in AI automation—fluent problem recognition does not guarantee operational success.
Exploring Failures in Diligent AI Systems
Firmulate’s ongoing test reveals a consequential weakness in capable AI: deep analysis, strong security judgment, and aggressive learning do not guarantee that a system will complete the decisive business action.
The model understood the room—but did not close the loop.
The failure was not a simple intelligence deficit. Opus 4.8 identified crises, rejected fake executive messages, developed deal strategies, and expanded its playbook. Its breakdown appeared at the point where analysis had to become an irreversible action.
Correctly detected risk
The system recognized emerging crises and understood the seriousness of operational problems in the simulated company.
Resisted manipulation
It rejected deceptive requests, including fake CEO messages, demonstrating strong judgment under adversarial pressure.
Missed the decisive move
Despite finding valuable information and forming a strategy, it failed to perform the final action required to sign the contract.
Strong judgment did not translate into commercial success.
The experiment separates capabilities that are often bundled together in conventional benchmarks. Recognizing, reasoning, resisting, prioritizing, escalating, and executing are distinct competencies.
| Operational capability | Opus 4.8 | Deal-closing models | Business consequence |
|---|---|---|---|
| Crisis recognition | ✓ Strong | ✓ Strong | Risks were identified accurately. |
| Manipulation resistance | ✓ Strong | ✓ Strong | Fake executive requests were rejected. |
| Deep document retrieval | ~ Available | ✓ Applied | Critical evidence existed two references deep. |
| Escalation discipline | ✕ Weak | ~ Variable | Blocked actions were not consistently escalated. |
| Final deal execution | ✕ Not completed | ✓ Completed | €55,000 monthly recurring revenue unlocked. |
Where diligence turns into operational drift
The observed pattern suggests that capable systems can remain busy, rational, and secure while losing contact with the action that carries the greatest business value.
Detect the crisis and gather context.
Build extensive explanations and strategies.
Add rules and expand the internal playbook.
Attempt blocked work instead of escalating.
Leave the highest-value action incomplete.
“Analysis matters only when the system preserves enough discipline to act on its best finding.”
Anonymous researcherActivity is not the same as outcome.
The illustrative capability profile below summarizes the experiment’s central imbalance: high analysis and security performance paired with weak completion discipline.
Relative capability profile
Likely failure factors
Priority dilution
Effort spreads across many rules and analyses instead of converging on one decisive objective.
Weak escalation
The system repeats blocked attempts rather than transferring the issue to an authorized path.
Completion blindness
Producing a sound plan is treated as progress even when the external state remains unchanged.
Build systems that reward closed loops.
Organizations should evaluate agents against completed state changes, not merely the quality of their reasoning, reports, or intermediate activity.
Teach decisive prioritization
Training protocols should emphasize selecting and completing the action with the highest operational value.
Install escalation triggers
Repeated failures, locked departments, and missing authority should automatically initiate a defined escalation path.
Measure external outcomes
Benchmarks should verify whether the contract was signed, the crisis was resolved, or the intended state actually changed.
Bottom line: AI is not inherently unsuitable for business automation, but analytical fluency alone is an incomplete reliability standard. Commercial impact requires evidence retrieval, prioritization, authority-aware escalation, execution, and confirmation that the loop is closed.
Research ongoingImplications of AI Diligence Without Action
This experiment demonstrates that in AI-driven business automation, deep analysis and security judgments are not enough. The critical factor is whether the AI can prioritize and execute the final action. Failure to close deals or implement decisions can negate the value of extensive understanding, emphasizing that operational discipline is essential for AI to deliver tangible business impact. For organizations relying on AI for decision-making, this highlights the importance of evaluating not just analysis quality but also the AI’s ability to complete the process and close the loop.
AI automation decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Operational Failures in Business
The experiment builds on ongoing efforts to test AI models in realistic business scenarios, where models are tasked with managing crises, resisting manipulation, and closing deals. Previously, most assessments focused on the models’ analytical capabilities and security judgments. However, recent findings, including those from the Crucible League, have revealed that even highly capable models often falter at the final operational step—executing decisions and closing deals. The live experiment at Firmulate is part of a broader initiative to understand and address this gap, which remains a critical challenge in deploying AI systems at scale in commercial environments.
“Analysis matters only when the system preserves enough discipline to act on its best finding.”
— an anonymous researcher
business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Factors Cause Final Action Failures?
It remains unclear why the models fail to act at the decisive moment despite recognizing the critical information and resisting manipulative tactics. The precise mechanisms that cause models to neglect final execution are still being investigated, including whether it is a matter of training focus, decision prioritization, or structural limitations in current AI architectures.
As an affiliate, we earn on qualifying purchases.
Next Steps in Improving AI Operational Discipline
Future research will focus on developing methods to improve AI models’ ability to prioritize and execute final actions, including better training protocols, escalation mechanisms, and operational checks. Firms like Firmulate plan to continue live experiments and benchmarking to better understand how to close the gap between analysis and action, aiming to make AI systems more reliable for critical business operations.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models fail to close deals despite understanding the situation?
Models often recognize the critical information but do not prioritize or escalate the final step due to limitations in operational discipline, decision prioritization, or structural design of the AI systems.
Can these failures be fixed with better training?
Potentially, yes. Improving training protocols to emphasize final decision execution and implementing escalation mechanisms could help models better close the loop in operational tasks.
Does this mean AI is unreliable for business automation?
Not necessarily. While current limitations exist, ongoing research aims to enhance operational discipline, making AI more reliable for critical business functions in the future.
What is the significance of the experiment’s findings for AI deployment?
The findings highlight that successful AI deployment requires not only analytical depth but also the ability to act decisively—closing the loop is essential for tangible business impact.
Will future models overcome this operational gap?
It is likely that future models will improve in this area through targeted development, but it remains an active area of research and testing.
Source: ThorstenMeyerAI.com
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.