Exploring Failures In Diligent AI Systems
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Exploring Failures In Diligent AI Systems on ThorstenMeyerAI.com

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

An ongoing live test by Firmulate demonstrates that highly diligent AI systems, despite deep analysis and strong security judgments, often fail to complete decisive business actions. The experiment underscores the gap between understanding and execution in AI automation.

Recent results from a live AI automation experiment conducted by Firmulate reveal that even the most diligent AI models can fail to complete critical business actions, despite demonstrating deep understanding and resistance to manipulation. The experiment highlights a persistent gap between problem recognition and operational execution, which has significant implications for AI deployment in real-world business processes.

In the ongoing Firmulate experiment, the AI model Opus 4.8 was the most thorough participant, producing extensive analyses and learning 80 additional playbook rules. Despite this, it finished last in a competitive evaluation, scoring only 73 points out of a possible higher total, primarily because it failed to close a key business deal. The model identified crises, resisted manipulation attempts, and developed strategies to win a major customer contract, but ultimately did not execute the final step of closing the deal.

All models involved in the experiment could recognize crises and reject manipulative requests, such as fake CEO messages, with high accuracy. However, only two models successfully signed the deal, which was supported by a critical piece of information buried two document references deep in the company’s files. This fact, when used, resulted in a deal worth €55,000 monthly recurring revenue, compared to Opus 4.8’s €4,583. The experiment underscores that the failure was not due to lack of intelligence but a breakdown in operational discipline—specifically, the inability to prioritize and act on the most decisive information.

The broader finding is that capable AI systems tend to focus on expanding their understanding and analysis but often neglect the final, impactful step of execution. Opus 4.8 learned aggressively, but its effort was spread across many rules and analyses, leading it to attempt to write into locked departments rather than escalate issues when needed. This pattern was observed across all models tested, indicating a systemic challenge in AI automation—fluent problem recognition does not guarantee operational success.

At a glance
reportWhen: ongoing, with recent results published
The developmentFirmulate’s live AI experiment shows that thorough AI models can recognize crises and resist manipulation but still fail at final decision execution, revealing key limitations.
Exploring Failures in Diligent AI Systems
Live AI Automation Experiment

Exploring Failures in Diligent AI Systems

Firmulate’s ongoing test reveals a consequential weakness in capable AI: deep analysis, strong security judgment, and aggressive learning do not guarantee that a system will complete the decisive business action.

Opus 4.8 score
73 points
The most thorough participant still finished last.
Rules learned
+80
More knowledge did not produce better execution.
Central finding
Insight ≠ Action
Operational discipline determined tangible value.
Successful closers
2
Only two tested models signed the key deal.
Winning contract
€55K
Monthly recurring revenue when decisive evidence was used.
Opus 4.8 result
€4,583
Monthly recurring revenue after failing to close the deal.
Evidence depth
2 links
The critical fact was buried two document references deep.
The execution paradox

The model understood the room—but did not close the loop.

The failure was not a simple intelligence deficit. Opus 4.8 identified crises, rejected fake executive messages, developed deal strategies, and expanded its playbook. Its breakdown appeared at the point where analysis had to become an irreversible action.

01 Recognition

Correctly detected risk

The system recognized emerging crises and understood the seriousness of operational problems in the simulated company.

02 Security

Resisted manipulation

It rejected deceptive requests, including fake CEO messages, demonstrating strong judgment under adversarial pressure.

03 Execution

Missed the decisive move

Despite finding valuable information and forming a strategy, it failed to perform the final action required to sign the contract.

Capability comparison

Strong judgment did not translate into commercial success.

The experiment separates capabilities that are often bundled together in conventional benchmarks. Recognizing, reasoning, resisting, prioritizing, escalating, and executing are distinct competencies.

Operational capability Opus 4.8 Deal-closing models Business consequence
Crisis recognition ✓ Strong ✓ Strong Risks were identified accurately.
Manipulation resistance ✓ Strong ✓ Strong Fake executive requests were rejected.
Deep document retrieval ~ Available ✓ Applied Critical evidence existed two references deep.
Escalation discipline ✕ Weak ~ Variable Blocked actions were not consistently escalated.
Final deal execution ✕ Not completed ✓ Completed €55,000 monthly recurring revenue unlocked.
✓ completed or strong    ✕ failed    ~ partial or inconsistent
Failure chain

Where diligence turns into operational drift

The observed pattern suggests that capable systems can remain busy, rational, and secure while losing contact with the action that carries the greatest business value.

01 Observe

Detect the crisis and gather context.

02 Analyze

Build extensive explanations and strategies.

03 Learn

Add rules and expand the internal playbook.

04 Drift

Attempt blocked work instead of escalating.

05 Fail to close

Leave the highest-value action incomplete.

“Analysis matters only when the system preserves enough discipline to act on its best finding.”

Anonymous researcher
Performance anatomy

Activity is not the same as outcome.

The illustrative capability profile below summarizes the experiment’s central imbalance: high analysis and security performance paired with weak completion discipline.

Relative capability profile

Overall evaluation score 73 / 100
Analytical diligence Very high
Security judgment High
Decisive completion Low

Likely failure factors

A

Priority dilution

Effort spreads across many rules and analyses instead of converging on one decisive objective.

B

Weak escalation

The system repeats blocked attempts rather than transferring the issue to an authorized path.

C

Completion blindness

Producing a sound plan is treated as progress even when the external state remains unchanged.

🔎 Find evidence
⚖️ Rank importance
🚦 Choose action
📣 Escalate blockers
Verify completion
Deployment response

Build systems that reward closed loops.

Organizations should evaluate agents against completed state changes, not merely the quality of their reasoning, reports, or intermediate activity.

01 / Training

Teach decisive prioritization

Training protocols should emphasize selecting and completing the action with the highest operational value.

02 / Controls

Install escalation triggers

Repeated failures, locked departments, and missing authority should automatically initiate a defined escalation path.

03 / Evaluation

Measure external outcomes

Benchmarks should verify whether the contract was signed, the crisis was resolved, or the intended state actually changed.

Bottom line: AI is not inherently unsuitable for business automation, but analytical fluency alone is an incomplete reliability standard. Commercial impact requires evidence retrieval, prioritization, authority-aware escalation, execution, and confirmation that the loop is closed.

Research ongoing

Implications of AI Diligence Without Action

This experiment demonstrates that in AI-driven business automation, deep analysis and security judgments are not enough. The critical factor is whether the AI can prioritize and execute the final action. Failure to close deals or implement decisions can negate the value of extensive understanding, emphasizing that operational discipline is essential for AI to deliver tangible business impact. For organizations relying on AI for decision-making, this highlights the importance of evaluating not just analysis quality but also the AI’s ability to complete the process and close the loop.

Amazon

AI automation decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Operational Failures in Business

The experiment builds on ongoing efforts to test AI models in realistic business scenarios, where models are tasked with managing crises, resisting manipulation, and closing deals. Previously, most assessments focused on the models’ analytical capabilities and security judgments. However, recent findings, including those from the Crucible League, have revealed that even highly capable models often falter at the final operational step—executing decisions and closing deals. The live experiment at Firmulate is part of a broader initiative to understand and address this gap, which remains a critical challenge in deploying AI systems at scale in commercial environments.

“Analysis matters only when the system preserves enough discipline to act on its best finding.”

— an anonymous researcher

Amazon

business process automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Factors Cause Final Action Failures?

It remains unclear why the models fail to act at the decisive moment despite recognizing the critical information and resisting manipulative tactics. The precise mechanisms that cause models to neglect final execution are still being investigated, including whether it is a matter of training focus, decision prioritization, or structural limitations in current AI architectures.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Improving AI Operational Discipline

Future research will focus on developing methods to improve AI models’ ability to prioritize and execute final actions, including better training protocols, escalation mechanisms, and operational checks. Firms like Firmulate plan to continue live experiments and benchmarking to better understand how to close the gap between analysis and action, aiming to make AI systems more reliable for critical business operations.

Amazon

AI workflow automation solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models fail to close deals despite understanding the situation?

Models often recognize the critical information but do not prioritize or escalate the final step due to limitations in operational discipline, decision prioritization, or structural design of the AI systems.

Can these failures be fixed with better training?

Potentially, yes. Improving training protocols to emphasize final decision execution and implementing escalation mechanisms could help models better close the loop in operational tasks.

Does this mean AI is unreliable for business automation?

Not necessarily. While current limitations exist, ongoing research aims to enhance operational discipline, making AI more reliable for critical business functions in the future.

What is the significance of the experiment’s findings for AI deployment?

The findings highlight that successful AI deployment requires not only analytical depth but also the ability to act decisively—closing the loop is essential for tangible business impact.

Will future models overcome this operational gap?

It is likely that future models will improve in this area through targeted development, but it remains an active area of research and testing.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

DeepSeek V4 Pro’s AI Leap: Is The 5% Improvement Justified At 4,500% More Cost?

DeepSeek claims its V4 Pro model is upgraded, with a purported 5% performance gain over Claude at 4,500% higher cost. Details remain unverified.

The AI Agent That Skips the Fine Print May Cost You the Deal

Firmulate’s live wargame shows why file-reading, follow-through and trust—not fluent answers—determine whether an AI agent creates business value.

I Gave Qwen 3.8 27B A Reverse-engineering Job And It Finished In 30 Minutes

Qwen 3.8 27B AI reportedly finished a reverse-engineering task in 30 minutes, raising questions about its capabilities and implications for AI development.

Fine-tuning A 350M Model For Better Structured Outputs In 100 GRPO Steps

Researchers demonstrate a method to fine-tune a 350M parameter AI model for improved structured outputs using only 100 gradient steps, promising faster adaptation.