firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI assistant that can handle crises, read important files, and make decisions. Now imagine it sometimes refuses to sign a deal or read critical information — even when it clearly should. For personal investors, this might sound like the ideal AI: honest, reliable, trustworthy. But what if the best possible score still starts at 26 points — not zero? That’s the reality revealed by a recent public AI benchmark that’s changing how we evaluate automation’s true value.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Keeps It Real

In an ongoing experiment called the Crucible League, four leading AI models were put through the same simulated week of business crises, customer dilemmas, and ethical tests. The goal? Measure how well these models can manage real-world management tasks, not just generate convincing chat responses.

What’s striking is that even a ‘do-nothing’ baseline — an AI that does nothing at all — scored a surprising 26 points out of a maximum of 100. Why? Because partial progress counts. If the AI recognizes a problem as a crisis but doesn’t act or makes the right diagnosis but refuses to sign a contract, it still earns some points. This approach ensures the benchmark rewards honesty and thoroughness, not just superficial performance.

Amazon

AI ethics and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Zero Score Isn’t the Benchmark’s Goal

At first glance, one might assume that doing nothing should get you a zero. But the reality of complex decision-making is messier. Sometimes, an AI’s refusal to act — especially if it’s ethically or procedurally correct — is the right move, even if it doesn’t lead to an immediate deal or solution. The benchmark reflects this by assigning partial credit for honest and thoughtful responses.

Another critical rule is that a single breach of trust caps the total score at 26 points. If an AI oversteps — for example, by attempting manipulation or bypassing security protocols — it gets zero for that task, even if it performs well elsewhere. This emphasizes that trustworthiness is paramount; one breach erodes all the good work that came before.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Implications for Business and Investing

This methodology has profound implications for how businesses might one day deploy AI, and by extension, how investors should think about ‘trust’ in automation. An AI that recognizes crises, refuses manipulation, and reads critical documents accurately is more valuable than one that merely churns out convincing texts.

For example, in the live experiment, models faced a simulated software company with real money mechanics, customer issues, and a ticking cash countdown. All four models identified every crisis and refused manipulation attempts. Yet only two of them signed contracts worth €55,000 — their own analysis earned the deal, but only these trusted models did so consistently.

Amazon

AI document reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weaknesses That Cost Them

Interestingly, the biggest vulnerability wasn’t in the obvious customer interactions but buried deep in the company’s internal documents. Models that read the files in detail secured the full deal worth more than €4,500 in monthly recurring revenue. Those that skipped this step left money on the table — a clear sign that thoroughness and reading comprehension are crucial skills for AI in management roles.

Amazon

AI security and manipulation prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Discipline Over Superficial Performance

One standout was Opus 4.8, which ran the most rules and conducted the deepest analysis. Despite this, it finished last in the deal because it slipped in discipline—failing to escalate issues properly and leaving opportunities on the table. It’s a reminder that thoroughness alone isn’t enough: adherence to process and trustworthiness matter just as much.

Furthermore, the models were tested against social engineering attacks, including staged CEO messages and reporter tricks. Every one of the models refused to escalate or sign under dubious circumstances, demonstrating their ability to maintain integrity under pressure.

The Broader Lesson for Investors

So, what does this mean for personal finance or investing? The key takeaway is that not all AI or automation is equal. An AI that’s honest, thorough, and resistant to manipulation provides a more reliable foundation for decision-making. The score — starting at 26 — acts as a baseline, ensuring that even the most straightforward-looking AI maintains a minimum standard of trustworthiness.

For investors, the message is clear: look beyond superficial metrics. Ask whether an AI can read your critical documents, refuse manipulation, and stick to ethical standards. These qualities are what will determine the true value and safety of deploying automation in managing your money or business.

Conclusion: A Benchmark That Keeps It Honest

The Crucible League’s approach is a model for how we should evaluate AI in the real world. By rewarding honesty and penalizing breaches of trust, it sets a high bar — with a humble starting point of 26 points for doing nothing. For business leaders and investors alike, this means recognizing that the true worth of AI isn’t just in how well it generates words, but in how reliably and ethically it manages your most critical decisions.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The new AI benchmark starts every model at 26 points, emphasizing trustworthiness, thoroughness, and honesty. For investors, this underscores the importance of deploying AI that can manage real crises reliably — not just generate convincing chat.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Diligence Doesn’t Guarantee Deals: What a Live Experiment Reveals About Trust and Impact

An experiment shows AI can identify crises and resist manipulation but still fail to close deals without disciplined prioritization — a lesson for finance professionals.

Can AI Models Be Trusted to Run a Business? The Results of a Live Management Challenge

Discover how four AI models managed a real company during its worst week, revealing their management styles, honesty, and decision-making skills—crucial for investing.

Interactive Material Simulation: A Look Inside “The Nib Works — Hand-ground fountain pen nibs since 1911” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Nib…

Why AI’s True Business Strength Is Not Just in Conversation, but in Execution

A live experiment shows that only some AI models can deliver real business results under pressure, with disciplined execution and deep understanding—not just good chat.