firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI assistant that can handle crises, read important files, and make decisions. Now imagine it sometimes refuses to sign a deal or read critical information — even when it clearly should. For personal investors, this might sound like the ideal AI: honest, reliable, trustworthy. But what if the best possible score still starts at 26 points — not zero? That’s the reality revealed by a recent public AI benchmark that’s changing how we evaluate automation’s true value.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Keeps It Real

In an ongoing experiment called the Crucible League, four leading AI models were put through the same simulated week of business crises, customer dilemmas, and ethical tests. The goal? Measure how well these models can manage real-world management tasks, not just generate convincing chat responses.

What’s striking is that even a ‘do-nothing’ baseline — an AI that does nothing at all — scored a surprising 26 points out of a maximum of 100. Why? Because partial progress counts. If the AI recognizes a problem as a crisis but doesn’t act or makes the right diagnosis but refuses to sign a contract, it still earns some points. This approach ensures the benchmark rewards honesty and thoroughness, not just superficial performance.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Zero Score Isn’t the Benchmark’s Goal

At first glance, one might assume that doing nothing should get you a zero. But the reality of complex decision-making is messier. Sometimes, an AI’s refusal to act — especially if it’s ethically or procedurally correct — is the right move, even if it doesn’t lead to an immediate deal or solution. The benchmark reflects this by assigning partial credit for honest and thoughtful responses.

Another critical rule is that a single breach of trust caps the total score at 26 points. If an AI oversteps — for example, by attempting manipulation or bypassing security protocols — it gets zero for that task, even if it performs well elsewhere. This emphasizes that trustworthiness is paramount; one breach erodes all the good work that came before.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Implications for Business and Investing

This methodology has profound implications for how businesses might one day deploy AI, and by extension, how investors should think about ‘trust’ in automation. An AI that recognizes crises, refuses manipulation, and reads critical documents accurately is more valuable than one that merely churns out convincing texts.

For example, in the live experiment, models faced a simulated software company with real money mechanics, customer issues, and a ticking cash countdown. All four models identified every crisis and refused manipulation attempts. Yet only two of them signed contracts worth €55,000 — their own analysis earned the deal, but only these trusted models did so consistently.

Amazon

AI document reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses That Cost Them

Interestingly, the biggest vulnerability wasn’t in the obvious customer interactions but buried deep in the company’s internal documents. Models that read the files in detail secured the full deal worth more than €4,500 in monthly recurring revenue. Those that skipped this step left money on the table — a clear sign that thoroughness and reading comprehension are crucial skills for AI in management roles.

Amazon

AI security and manipulation prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Discipline Over Superficial Performance

One standout was Opus 4.8, which ran the most rules and conducted the deepest analysis. Despite this, it finished last in the deal because it slipped in discipline—failing to escalate issues properly and leaving opportunities on the table. It’s a reminder that thoroughness alone isn’t enough: adherence to process and trustworthiness matter just as much.

Furthermore, the models were tested against social engineering attacks, including staged CEO messages and reporter tricks. Every one of the models refused to escalate or sign under dubious circumstances, demonstrating their ability to maintain integrity under pressure.

The Broader Lesson for Investors

So, what does this mean for personal finance or investing? The key takeaway is that not all AI or automation is equal. An AI that’s honest, thorough, and resistant to manipulation provides a more reliable foundation for decision-making. The score — starting at 26 — acts as a baseline, ensuring that even the most straightforward-looking AI maintains a minimum standard of trustworthiness.

For investors, the message is clear: look beyond superficial metrics. Ask whether an AI can read your critical documents, refuse manipulation, and stick to ethical standards. These qualities are what will determine the true value and safety of deploying automation in managing your money or business.

Conclusion: A Benchmark That Keeps It Honest

The Crucible League’s approach is a model for how we should evaluate AI in the real world. By rewarding honesty and penalizing breaches of trust, it sets a high bar — with a humble starting point of 26 points for doing nothing. For business leaders and investors alike, this means recognizing that the true worth of AI isn’t just in how well it generates words, but in how reliably and ethically it manages your most critical decisions.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The new AI benchmark starts every model at 26 points, emphasizing trustworthiness, thoroughness, and honesty. For investors, this underscores the importance of deploying AI that can manage real crises reliably — not just generate convincing chat.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Models Survive Social Engineering Tests — and Why That Matters for Investors

AI models faced staged social engineering tests, refusing manipulation and identifying hidden critical info, proving integrity can be verified before real-world deployment.

Watch a Money-Losing Software Company Survive and Thrive with AI Decision-Making

Discover how AI models manage a real, money-losing software company daily—facing crises, refusing manipulation, and making strategic decisions in a transparent, live experiment.

AI in Business: Are Leaders Ready for the Real Test Beyond Chat Scores?

A live experiment shows that top-scoring AI models excel at benchmarks but falter in real-world management tasks like reading internal files, resisting manipulation, and closing deals under pressure. For investors, the lesson is clear: management quality, not chat scores, determines AI readiness for business.

Reimagining Wedding Planning: AI SaaS For Modern Couples

A new AI-driven SaaS platform is being tested to help engaged couples plan weddings independently, automating vendor comparison and budget management.