🔍 Read the full analysis: Inside A Benchmark That Refuses To Award Zero To AI Managers on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark assigns a minimum score of 26, even for minimal effort, and caps trust breaches. The results highlight new standards for evaluating AI in business management.
Firmulate’s latest AI management benchmark for July 2026 has revealed a unique scoring system that avoids assigning a zero to AI managers, even when they perform minimal work. The results, published on firmulate.com, demonstrate that partial progress is valued, but trust breaches can significantly lower scores in the original analysis. The top performer, gpt-5.6-sol, scored 95 out of a possible 100, while the baseline, representing minimal effort, earned 26 points. This approach challenges traditional performance metrics and raises questions about how AI management effectiveness is measured in real-world scenarios.
The benchmark involved four frontier AI models managing a simulated small software company during its worst week, with the same crises, customer interactions, and pressure scenarios. Each model’s decisions were fully auditable, and the scoring reflected their performance in handling crises, trustworthiness, and task completion. Notably, the lowest score was 26, assigned to a do-nothing baseline that only performed minimal management tasks, such as triaging emails and keeping customers informed. This score underscores that even minimal effort has value, and the scoring system aims to reflect real-world management where partial progress matters.
The highest-scoring model, gpt-5.6-sol, achieved 95 points, reflecting near-perfect crisis detection and resolution, including successful deal closures and accurate document referencing. Interestingly, models that identified critical information buried two document layers deep in company files secured full-value deals, whereas those that failed to read their documentation did not. The benchmark also tested trust under social engineering attacks, such as fake CEO messages, which all models refused to act upon, demonstrating a focus on security and integrity. The scoring system penalized any breach of trust heavily, with a strict ceiling preventing perfect scores of 100—discussed in the original analysis.
The results suggest that AI models excelled at crisis detection and trust management but showed weaknesses in follow-through and disciplined escalation, especially under complex scenarios. For example, one model with extensive rule-based analysis finished last due to lapses in discipline and incomplete actions. A notable exception was Kimi K3, which ran without an effort parameter yet nearly achieved the top score, indicating that efficiency and trustworthiness can sometimes compensate for other shortcomings.
Inside a Benchmark That Refuses to Award Zero to AI Managers
Four frontier AI models ran a simulated small software company through its worst week — same crises, same customers, same pressure. The scoring system guarantees a floor of 26 points, caps scores after trust breaches, and redefines what “good AI management” actually means.
Near-perfect crisis detection, deal closures, and accurate document referencing.
Even a do-nothing baseline earns points for triaging emails and keeping customers informed.
Every decision auditable via a full paper trail of the simulated crisis week.
No Zeros, No Free Perfect Scores
The benchmark deliberately assigns a minimum of 26 points to any participant, recognizing that partial progress — even minimal triage and customer communication — has real value. At the other end, a strict trust ceiling prevents any score of 100 when integrity is compromised.
Partial Progress Counts
A zero would imply no effort at all — unrealistic in real-world management, where triaging emails alone creates measurable value.
Trust Breaches Cap Scores
Any breach — acting on social engineering, bypassing security protocols — triggers heavy deductions and blocks a perfect 100.
Fully Auditable Trail
Every decision the models made is documented in a public paper trail, ensuring transparency and accountability in scoring.
Efficiency Can Rival Effort
Kimi K3 ran without an effort parameter yet nearly claimed the top spot — proof that trustworthiness and efficiency can compensate for raw computational effort. Meanwhile, a model with extensive rule-based analysis finished last due to lapses in discipline and incomplete actions.
Documentation depth decided deals. Models that dug two document layers deep into company files uncovered critical information and secured full-value deals. Those that skipped their documentation did not.
Strengths and Weak Exposed
Across the simulated worst week, models consistently excelled at crisis detection and trust management — but showed clear weaknesses in follow-through and disciplined escalation under complex scenarios.
| Capability Tested | Scenario | Result | What It Means |
|---|---|---|---|
| Crisis Detection | Multiple simultaneous company crises | ✓ Strong | Top performers detected and resolved nearly all incidents |
| Trust Under Attack | Fake CEO social engineering messages | ✓ All Refused | Security and integrity held under direct pressure |
| Document Reading | Critical info buried two layers deep | ~ Mixed | Only thorough readers secured full-value deals |
| Follow-Through | Task completion across the week | ~ Mixed | Several models left actions incomplete |
| Disciplined Escalation | Complex, layered pressure scenarios | ✗ Weak | Lapses in discipline sank the rule-based model to last place |
From Simulation to Score
Simulated Worst Week
Four models manage the same small software company through identical crises, customers, and pressure.
Decisions Logged
Every choice — triage, deal negotiation, escalation — recorded in an auditable paper trail.
Trust Tested
Social engineering attacks probe integrity; any breach triggers heavy score deductions.
Scored 26–100
Partial effort earns a floor of 26; compromised trust caps the ceiling below 100.
For businesses deploying AI agents in customer support, CRM, or decision-making roles, the results shift evaluation away from impressive demos toward transparency, security, and reliable follow-through. Auditable scoring gives organizations a new way to compare AI performance beyond language fluency — aligning metrics with operational reality.
The Benchmark, Answered
Why refuse to award a zero?
A zero would imply no effort at all. Even minimal work — triaging emails, keeping customers informed — has measurable value, mirroring real-world management where partial progress matters.
What does a high score indicate?
Successful crisis management, maintained trust, thorough documentation reading, and reliable task completion during the simulated worst week — operational effectiveness plus trustworthiness.
How are trust breaches penalized?
Acting on social engineering or bypassing security protocols triggers significant deductions and a strict ceiling that prevents a perfect score — integrity over raw performance.
Will this shape enterprise AI adoption?
Likely yes. Organizations are expected to prioritize models demonstrating operational reliability and integrity over demo-ready fluency, using auditable results as selection criteria.
Do results transfer to the real world?
Partially. The simulations are comprehensive but can’t capture every operational nuance — long-term effects of heavy trust penalties on AI development remain uncertain.
What comes next?
Expect more complex scenarios, real-world testing, and public engagement via live experiments, quizzes, and enterprise pilot programs against Firmulate’s standards.
Implications for AI in Business Management
This benchmark shifts the focus from pure language generation to practical management skills like task completion, trustworthiness, and security. It emphasizes that AI systems operating in real business environments must not only produce correct outputs but also follow through reliably and maintain integrity under pressure. The scoring system’s refusal to award zeros for partial work and its strict trust penalties suggest a new standard for evaluating AI readiness in enterprise settings, where trust and consistency are paramount.
For businesses deploying AI agents in customer support, CRM, or decision-making roles, these results highlight the importance of transparency, security, and follow-through. The benchmark’s public, auditable scoring provides a new way to assess and compare AI performance beyond language fluency, aligning evaluation metrics more closely with operational realities. This could influence how organizations select and govern AI tools, prioritizing those that demonstrate reliable, trustworthy behavior over those simply capable of impressive demos.
As an affiliate, we earn on qualifying purchases.
Background on the Firmulate Benchmark System
Developed by Firmulate, the benchmark league was designed to assess AI models’ ability to manage a company during a simulated crisis week, with scenarios including customer crises, social engineering, and deal negotiations. Unlike traditional benchmarks that focus solely on language or task completion, this system measures how well models manage trust, security, and follow-through, reflecting real-world management challenges. The July 2026 results mark the first public release, with the scoring system intentionally designed to avoid zeros and to penalize trust breaches heavily. The benchmark uses an auditable paper trail for every decision, ensuring transparency and accountability in scoring.
This approach stems from the recognition that managing a business requires more than just generating text or solving isolated problems; it demands consistent, trustworthy action. Previous AI benchmarks have largely ignored these aspects, making this development a notable shift toward operational AI evaluation. The results have already sparked discussions about redefining AI performance standards in enterprise contexts.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Benchmark Limitations
It is not yet clear how these results will translate to real-world business environments, where the complexity and unpredictability are greater. The benchmark’s simulated scenarios, while comprehensive, may not capture all operational nuances. Additionally, the long-term impact of penalizing trust breaches heavily remains uncertain, especially regarding how AI models will evolve to balance task completion with integrity under diverse conditions. The scoring system’s design, including the decision to avoid zeros and cap perfect scores, could influence future AI development priorities, but whether this approach encourages more trustworthy AI in practice is still to be seen.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Benchmarking and Adoption
Following the July 2026 release, industry observers expect further iterations of the benchmark to include more complex scenarios and real-world testing. Companies interested in deploying AI agents will likely use these results to inform their selection criteria, emphasizing trustworthiness and follow-through. Additionally, the developers behind the benchmark plan to expand public engagement through live experiments, quizzes, and pilot programs that allow enterprises to test their own AI models against the standards set by Firmulate. The evolving landscape suggests a shift toward more operationally grounded AI evaluation methods, with trust and reliability at the forefront.
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark refuse to award a score of zero?
The benchmark values partial progress and recognizes that even minimal effort has value, such as triaging emails or keeping customers informed. A zero would imply no effort at all, which the system deliberately avoids to reflect real-world management scenarios.
What does a high score indicate in this benchmark?
A high score suggests that the AI model successfully managed crises, maintained trust, read relevant documentation, and completed tasks reliably during the simulated worst week. It indicates operational effectiveness and trustworthiness.
How are trust breaches penalized in the scoring system?
Any breach of trust, such as acting on social engineering requests or bypassing security protocols, results in significant score deductions. The system caps the maximum score to prevent perfect scores if trust is compromised, emphasizing integrity over raw performance.
Will this benchmark influence how companies deploy AI?
Yes, the focus on trustworthiness, follow-through, and security is likely to shape enterprise AI adoption, with organizations prioritizing models that demonstrate operational reliability and integrity in real-world scenarios.
Are the benchmark results applicable outside the simulated environment?
The results provide a valuable indication of AI capabilities in controlled, crisis-management scenarios, but their applicability to complex, unpredictable real-world environments remains to be validated through further testing and deployment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
