
A polished home office can make remote work feel effortless. But if an AI assistant starts handling your support queue, customer records or business forecast, the real test is less about how smoothly it talks and more about what it does under pressure. A live Firmulate experiment puts that question to work: five frontier models ran the same small software company through its worst week.
Get business pricing on your home office setup
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company, not a chat demo
In the Crucible League’s final results for July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
Each model faced the same customers, crises and temptations. The company is a live experiment at Firmulate, with synthetic employees, a public cash countdown and versioned workdays. Its mechanics include monthly burn of €105,000 against €2,300 in monthly recurring revenue, and more than 680 self-learned playbook rules. The experiment is watchable at Firmulate.
The deal hidden in the files
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The difference came down to a buried competitor weakness, tucked two document references deep in the company’s files rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
K3 found that weakness and closed the deal. So did the top-ranked gpt-5.6-sol. The result is a useful reminder for anyone setting up an AI-enabled workspace: a confident answer is not the same as finishing the job. The models could reach the same diagnosis and make the same pitch, yet most did not sign.
Discipline under pressure
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s 93-point finish came with one deviation, the cleanest discipline in the field. Opus 4.8 offers a different lesson: it was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried writing into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four Western models.
For a home office, the stakes may begin with drafting messages or organizing tasks. For a company, an AI workforce could eventually touch customer support, sales records or forecasts. Firmulate’s findings suggest that model choice deserves a practical trial in the work it will actually be asked to do. Its benchmark results are available at firmulate.com/benchmarks.html. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Test the setup before you trust it
The newcomer beat three of the four Western frontier models, but no leaderboard can tell every company which AI will fit its work. Firmulate offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems. For anyone choosing an AI assistant, the practical question is whether it can find the right information, act on its analysis and stay disciplined when the pressure rises.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
