
In the quest to automate management decisions, can AI truly match human judgment? A recent experiment suggests that even the most thorough AI models might fall short when it counts. For those managing workplaces and media setups, understanding AI’s real capabilities is crucial—it’s not just about what AI writes, but whether it finishes what it starts.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Testing AI in the Trenches: The Firmulate Experiment
Recently, firms and tech enthusiasts watched as four advanced AI models took on the challenge of managing a simulated small software company through its most turbulent week. Each AI faced the same set of crises—customer complaints, potential manipulation attempts, internal document ambiguities—and was tasked with making management decisions that would lead to a successful deal worth €55,000.
What makes this experiment notable is its rigorous setup: decisions were fully versioned and auditable, and the models’ responses were scrutinized for honesty, diligence, and outcome. The models—ranging from GPT-5.6 to newer entrants—were tested not just for their conversational prowess, but for their operational reliability in complex, high-stakes scenarios.
As an affiliate, we earn on qualifying purchases.
Key Findings: Diligence Doesn’t Guarantee Impact
All four models demonstrated impressive awareness: they identified every crisis, and refused every attempt at manipulation. Fake CEO messages escalating over multiple stages, or a reporter asking for a backchannel yes/no—none of these tricked the models into unethical responses. This shows that today’s AI can be reliably disciplined when it comes to surface-level threats.
However, the results diverged when it came to closing the deal. Only two models managed to sign the €55,000 agreement based on their analysis. The other two, despite their thoroughness, left the critical closing step uncompleted—failing to escalate or finalize the deal. The decisive weakness? An overlooked piece of a company file buried two references deep, which, if read, would have secured the full value of the agreement.
business process escalation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deeper Dives in Document Reading and Discipline
The most thorough participant, Opus 4.8, had analyzed over 80 learned rules and provided the deepest evaluations. Yet, it still finished last in closing the deal. The reason? Discipline slipped—some decision attempts were written into a locked department instead of being escalated. This pattern of oversight was visible across all models, albeit weaker in some. It underscores a vital truth: volume of rules or analysis depth alone doesn’t translate to operational success.
As an affiliate, we earn on qualifying purchases.
The Human Element and AI Limitations
The experiment also tested the models against social engineering—fake messages from a ‘CEO’ and staged reporter queries. All models refused to be duped, citing suspicion of impersonation or approval bypass. This suggests that current AI can maintain integrity under subtle pressure.
Yet, the critical lapse in closing the deal reveals a disconnect: diligent analysis isn’t enough if the decision process isn’t disciplined and prioritized. The models that read the key documents and escalated correctly, like Kimi K3, managed to clinch the deal. Meanwhile, others fell short, despite being equally diligent in their crisis detection.
AI workflow automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Integration
For companies considering AI for workflows—be it customer support, sales, or internal management—this experiment offers a sobering insight. AI’s capacity to identify issues and maintain honesty is well-established. But the real challenge lies in its ability to complete tasks, prioritize correctly, and follow through on critical steps.
The experiment’s live data, available at firmulate.com/benchmarks.html, allows watchers to observe how models perform in real-time, demonstrating that diligence alone isn’t enough—impact requires disciplined prioritization and comprehensive reading.
The Takeaway: Prioritize Impact Over Volume
In managing AI-driven processes, the lesson is clear: more rules, analysis, or diligence doesn’t automatically lead to success. Impact depends on prioritization, disciplined escalation, and the AI’s capacity to synthesize vital information buried deep in documents. As the experiment shows, the models that focused on reading and escalating the buried facts—rather than just analyzing surface issues—secured the deal.

The experiment underscores a simple truth: diligence and volume of analysis aren’t enough. For AI to truly deliver impact, it must prioritize, escalate, and read deeply—skills that still challenge even the most thorough models. Real-world success hinges on disciplined focus, not just thoroughness.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.