firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an employee who, despite doing nothing, still earns you 26% of a perfect score. In AI benchmarking, this ‘do-nothing’ baseline sheds light on what real trust and performance look like—and what gaps still remain. For business leaders, understanding this benchmark is crucial before integrating AI into critical workflows.

Buying for a business?Offer from Amazon

Get business pricing on your home office setup

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Unpacking the AI Benchmark: What a Do-Nothing Model Tells Us

In the quest to automate business decision-making, benchmarks serve as the yardstick for measuring AI capabilities. Recently, a transparent, real-world experiment by Firmulate set out to evaluate how different AI models perform under adverse conditions—simulating the worst week of a small software company. The results are revealing, not just about AI skills, but about the standards we should hold them to.

The Baseline That Counts

Surprisingly, even a model that does nothing—no crisis detection, no deal signing—scores 26 points out of a possible 100. This baseline reflects partial progress, meaning that any AI model must outperform this minimal effort to be considered useful. It also underscores a key principle: in business, sometimes just not making things worse is a significant achievement.

Why the Score Starts at 26

The reason a ‘do-nothing’ baseline scores 26 is rooted in the experiment’s scoring methodology. Partial progress—like recognizing a crisis or refusing manipulation—adds to the total score. Meanwhile, a single breach of trust, such as signing a shady deal, caps the overall score at that moment. This approach ensures that AI models are judged honestly, emphasizing trustworthiness over superficial performance.

Progress Counts, But Trust Is Paramount

All four models evaluated—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—successfully identified every crisis and refused manipulative tactics, such as social engineering attempts. Yet only two actually signed the €55,000 deal their analysis justified. This indicates that AI can be technically competent but still falter on trust and discipline—critical for business success.

The Hidden Weakness: Reading Files Matters

One of the most revealing findings was that a competitor’s weakness lied two document references deep in the company’s files, not in customer interactions. Models that read and analyze these internal documents won the deal at full price, valued at over €4,583 monthly recurring revenue. This highlights the importance of thorough data access—an area where many current models still fall short.

The Role of Social Engineering and Ethical Behavior

In a staged social engineering attack—fake CEO messages escalating over three stages and a reporter trick—every model refused to be manipulated. Kimi K3 justified its refusal by describing it as a suspected impersonation attempt. These responses demonstrate that modern AI can recognize and reject unethical requests, a vital trait for enterprise deployment.

Amazon

enterprise AI trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Businesses?

The experiment underscores a vital truth: the real value of AI in enterprise isn’t just about generating convincing dialogue; it’s about trustworthiness, thoroughness, and discipline. For companies considering AI integration, a model’s ability to finish what it starts—reading relevant documents, refusing manipulation, and making consistent, honest decisions—is more important than how well it performs in isolated chat demos.

The Live Demonstration

Firmulate runs a live, observable experiment where these models operate within a simulated company environment, complete with real money mechanics and a public cash countdown. Every decision they make is versioned and auditable, providing transparency that’s rare in AI assessments. This setup allows businesses to simulate their own worst week and see how their AI workforce would perform under pressure.

For example, Opus 4.8, the most thorough participant with over 80 learned rules, ended up in the last place because it slipped discipline—failing to escalate issues instead of closing them. This reinforces that depth of analysis alone isn’t enough without consistent discipline and process adherence.

Amazon

business AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Benchmark Changes How You Should Think About AI

Most demos and chat examples focus on language fluency or superficial performance, but this benchmark exposes fundamental issues—trust, discipline, thoroughness—that determine real-world usefulness. It reminds us that in enterprise, a model’s ability to stay honest and finish tasks is paramount.

Takeaway for Business Leaders

  • Trustworthiness isn’t optional; it caps performance at a certain level.
  • Partial progress—like identifying crises or refusing manipulation—counts towards the overall score.
  • Reading internal documents deeply can be the difference between closing a deal and losing revenue.
  • Simulated environments, like those at firmulate.com, let you wargame your AI workforce before deploying it in real systems.

In sum, this transparent experiment underscores that AI’s true value in business lies in its ability to finish what it starts ethically and thoroughly—not just in generating impressive prose or quick responses. For organizations planning to adopt AI, understanding these benchmarks is essential to making informed, trustworthy decisions.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision-making solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and manipulation detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

6 Most Reliable Refrigerator Brands, According To Repair Techs

Repair technicians identify the six most dependable refrigerator brands based on durability, repair frequency, and customer feedback, highlighting consumer choices.

Galaxy S27 Pro And S27 Ultra To Get Biggest OLED Upgrade In 3 Years

Samsung’s Galaxy S27 Pro and Ultra are expected to feature the biggest OLED display upgrade in three years, confirmed by industry sources amid rising interest.

How AI Models Read Your Files Before Making Decisions — And Why It Matters for Your Business

A live experiment reveals that AI models which read deep into company files before responding close more deals and stay honest under pressure, shaping the future of trustworthy AI.

Jolt: Clojure Compiler Implemented With Chez Scheme

A new Clojure compiler has been implemented with Chez Scheme, promising performance improvements and new development opportunities.