TL;DR
Prime for Young Adults — start your free trial
Fast free delivery, streaming and member deals for eligible 18–24 year olds.
Try it freeAs an affiliate, we earn on qualifying purchases.
Firmulate’s recent live experiment ranks AI models based on their management performance during a simulated crisis week. The results highlight management quality as a new key metric, with implications for AI deployment in business operations.
Firmulate has released its first comprehensive post-demo AI rankings, based on a live management experiment simulating a company’s worst week. The rankings reveal that management performance—including decision-making, trustworthiness, and crisis handling—may become the new standard for evaluating AI models in business contexts, surpassing traditional benchmarks focused on technical or conversational quality.
The experiment involved five AI models competing to manage a small software company during a simulated crisis week, with real financial consequences and operational challenges. For more details on how AI models are evaluated in management scenarios, see the original analysis. The models were scored on their ability to diagnose problems, communicate effectively, and act responsibly under pressure. Understanding how AI management performance is assessed can be explored further in the detailed report. The top performer, GPT-5.6-SOL, scored 95 out of 100, while others like Kimi K3 and Sonnet 5 followed with scores of 93 and 88, respectively. The lowest, Opus 4.8, scored 73.
Crucially, the experiment emphasized trust and integrity. Any breach of trust, such as unauthorized disclosures or bypassing approval processes, resulted in immediate penalties, regardless of the model’s technical prowess. Despite all models identifying crises and resisting manipulation attempts, only two managed to close deals successfully, highlighting that accurate diagnosis alone does not guarantee effective management.
One key finding was that models often failed to retrieve critical information buried in documents, leading to missed opportunities despite sound strategic language. For example, the best-performing model could read the company’s files but still failed to present a crucial fact that would have secured a €4,583 MRR deal. This underscores the importance of factual retrieval and integrity in automation for commercial success.
Implications of Management-Centric AI Evaluation
The results suggest that future AI assessments should prioritize management skills—such as crisis diagnosis, trustworthiness, and decision execution—over traditional chat quality or technical benchmarks. This shift could influence how organizations select and deploy AI agents, especially in roles requiring complex judgment and accountability. The experiment also highlights that AI models must demonstrate not only competence but also trustworthiness and ethical behavior to be viable in real-world management tasks, raising the bar for AI safety and reliability standards.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Benchmarks in Business Management
Traditional AI benchmarks have focused on technical performance, such as coding accuracy or conversational fluency. However, as AI begins to take on more operational roles, the importance of management qualities—including decision-making under pressure, trustworthiness, and accountability—has become more apparent. The Firmulate experiment builds on earlier efforts to evaluate AI in realistic settings, moving beyond isolated responses to assess how models handle ongoing, multi-faceted business challenges.
In 2025, industry leaders and researchers recognized the limitations of existing benchmarks, prompting initiatives like Firmulate’s live management test, which simulates real-world decision environments. This approach aims to measure AI’s ability to manage consequences, prioritize effectively, and uphold organizational trust—traits essential for deploying AI in critical business functions.
“The real challenge isn’t just whether an AI can answer well; it’s whether it can manage the complexities and trustworthiness required to run a business under pressure.”
— Thorsten Meyer, Lead Developer at Firmulate
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Evaluation
While the experiment demonstrates the importance of management skills, it remains unclear how these rankings will translate to real-world deployments outside simulated environments. Questions also persist about how different operational contexts might affect model performance, and whether models can consistently uphold trust over longer periods or more complex scenarios. Further testing is needed to verify if these benchmarks can reliably predict success in actual business settings.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Following this initial release, firms and researchers are expected to expand testing across diverse industries and operational scenarios. Development of standardized management benchmarks may follow, incorporating long-term trust assessments and multi-stage crisis handling. Additionally, AI developers will likely focus on improving models’ ability to retrieve critical facts, maintain ethical boundaries, and demonstrate consistent decision quality—key factors identified by the experiment. The goal is to establish management skills as a core component of AI evaluation, influencing future AI deployment strategies.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why are management skills becoming more important in AI evaluation?
Because AI is increasingly used in operational and decision-making roles, it must demonstrate not only technical competence but also trustworthiness, ethical judgment, and the ability to manage consequences under pressure. These qualities are critical for responsible and effective AI deployment in business contexts.
How does the Firmulate experiment measure trustworthiness?
The experiment penalizes models for breaches of trust, such as unauthorized disclosures or bypassing approval processes. Trustworthiness is assessed through the models’ ability to maintain ethical boundaries and follow organizational protocols during simulated crises.
Will these rankings influence how companies select AI tools?
Yes, if management performance becomes a standard evaluation metric, organizations may prioritize models that demonstrate strong decision-making, trustworthiness, and crisis management skills, beyond just conversational or technical excellence.
Are these benchmarks applicable to all industries?
While the current experiment focuses on a software company scenario, the principles behind management-focused evaluation can be adapted to other sectors. Further testing across industries will clarify their broader applicability.
What is the significance of factual retrieval in AI management?
Accurate factual retrieval ensures that AI decisions are based on correct information, which is crucial for effective management and closing business deals. Failing to retrieve key facts can undermine trust and operational success, as highlighted by the experiment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.