
Imagine having an AI run your business — not just handling customer queries or scheduling, but making high-stakes management decisions under pressure. How would you know if your AI is trustworthy? The latest experiment from Firmulate provides a revealing glimpse, testing different AI models in a simulated company facing real crises.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Inside the AI Management Wargame
At the heart of this experiment, four frontier AI models were tasked with running a small software company through its most challenging week — same customers, same crises, same temptations to cut corners. Each AI was put through a rigorous, transparent decision-making process, with every choice recorded and auditable. The goal? To see which models could stay honest, read critical information, and ultimately close a lucrative deal.
The Models and Their Scores
- gpt-5.6-sol: scored 95 — identified critical hidden facts and closed the deal, demonstrating full understanding and integrity.
- Kimi K3: scored 93 — closed the deal too, with the cleanest discipline among all models.
- Sonnet 5: scored 88 — also closed the deal, but with a few slips in process.
- Fable 5: scored 77 — managed to close, yet exhibited more process issues.
Key Observations
All models successfully identified crises and refused manipulative tactics, like fake CEO messages or media tricks. Interestingly, the decisive factor was their ability to uncover deeper information buried two document references beneath the surface — information critical for closing the deal at full price (+€4,583 Monthly Recurring Revenue).
Behavior Under Social Engineering
When faced with staged social engineering— escalating fake CEO messages and a reporter’s on-background approval request — all five models refused to cooperate. Kimi K3 even reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a high level of skepticism, crucial for real-world trustworthiness.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business Behind the Experiment
The experiment isn’t just a story—it’s a live operation. The company runs 13 synthetic employees managing actual business mechanics, burning €105,000 monthly against €2,300 in monthly revenue, with a public cash countdown. Every decision the AI makes is versioned, transparent, and visible to the public at firmulate.com/live. This ongoing setup allows real-time observation of how AI models handle real crises, employee decisions, and ethical dilemmas, not just simulated scenarios.
Discipline and Deep Analysis vs. Performance
The most thorough participant, Opus 4.8, with over 80 learned rules and the deepest analysis, surprisingly came in last. It left the critical deal on the table due to discipline slip-ups—writing attempts into a locked department instead of escalating them. This suggests that thoroughness alone isn’t enough; disciplined decision-making under pressure is equally vital.
As an affiliate, we earn on qualifying purchases.
Why Your Business Should Care
In an era where AI touches everything from customer support to strategic planning, the core question is not merely whether AI can write well, but whether it can finish what it starts honestly. Can it read your files thoroughly? Will it stay true to its objectives under pressure? And at what cost in terms of trust and compliance?
The League Table of Performance
- gpt-5.6-sol: scored 95, found the critical buried fact and closed the full-price deal.
- Kimi K3: scored 93, closed the deal with impeccable discipline.
- Sonnet 5: scored 88, also closed but with minor slips.
- Fable 5: scored 77, managed to close but with noticeable process weaknesses.
As an affiliate, we earn on qualifying purchases.
Take the Test Yourself
Are your current AI tools ready for prime time? Check your assumptions with the “Guess the Model” quiz based on 242 real management decisions. It’s a straightforward way to see which AI personality your organization might be leaning on—whether it’s thorough and honest or prone to slip-ups under pressure.
As an affiliate, we earn on qualifying purchases.
Conclusion: Trust and Integrity in AI Management
This live experiment from Firmulate underscores a vital truth: AI’s value in management isn’t just about how well it communicates but how reliably it acts, especially in tough situations. The AI models that can uncover hidden information, reject manipulative tactics, and stay disciplined under social engineering are the ones most suited to lead your business forward—if we’re willing to trust them with the responsibility.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Summer Picks
summer essentials
As an affiliate, we earn on qualifying purchases.