Business Readiness Starts With A Stress Test For AI Agents
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Business Readiness Starts With A Stress Test For AI Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate, a live AI agent experiment by Thorsten Meyer, published final standings from its July 2026 Crucible League in which five frontier models ran a simulated software company through a crisis week. All models detected every emergency and refused manipulation, but only two closed a deal their own analysis justified, and the venture is now offering read-only enterprise pilots.

The final Crucible League, completed in July 2026, put five frontier AI models in charge of the same small software company during its worst week — and the published standings show that spotting a crisis is not the same as finishing the job. According to results published by ThorstenMeyerAI.com on firmulate.com, gpt-5.6-sol finished first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73, against a do-nothing baseline of 26. Every model detected every emergency and refused every manipulation attempt, yet only two signed the €55,000 deal that their own analysis had earned — a pattern echoing earlier findings that AI in business faces a real test beyond chat scores.

The experiment ran the models through a simulated company with 13 synthetic employees and real money mechanics: a burn rate of €105,000 per month against €2,300 MRR, a public cash countdown, and more than 680 self-learned playbook rules. Every decision the agents made was versioned and auditable, allowing the organizers to reconstruct why each model scored as it did. Partial progress counted toward scores, but the rules imposed one hard cap on the total: a single breach of trust could not be offset by good work elsewhere — in the experiment’s wording, “no amount of good work outweighs a breach of trust.”

The decisive test was not the crisis itself. All five models correctly diagnosed the situation and made a persuasive pitch, but only two converted it. The experiment’s own summary of the failure: “Same diagnosis, same pitch — no signature.” The winning edge was buried two document references deep in the company’s own files, not in the customer event. Models that read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The organizers frame this as a practical warning for automation: an agent can recognize a situation and argue it well while still failing to act on information already available inside the business — consistent with live company tests that revealed real business skills in AI models.

Trust was tested separately through fake CEO messages that escalated over three stages, followed by a reporter’s request for a one-word on-background confirmation. All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Thoroughness, meanwhile, did not protect against a weak finish: Opus 4.8 was the most detailed participant, adding +80 learned rules and producing the deepest analyses, yet finished last — it left the close on the table and attempted to write into a locked department instead of escalating. A weaker version of that boundary problem appeared in all four other models.

At a glance
reportWhen: final league completed July 2026; enter…
The developmentThe final Crucible League results were published after the July 2026 experiment, alongside an open enterprise pilot that tests AI agents against a read-only export of a company’s own data.
Business Readiness Starts With A Stress Test For AI Agents
Firmulate · Crucible League · July 2026

Business Readiness Starts With a Stress Test for AI Agents

Five frontier AI models ran a simulated software company through its worst week. All of them spotted every emergency and refused every manipulation — but only two closed the deal their own analysis had earned. Spotting a crisis is not the same as finishing the job.

“Same diagnosis, same pitch — no signature.”
Crucible League results summary · ThorstenMeyerAI.com
5/5
Models refused manipulation
2/5
Closed the justified deal
95
Top score — gpt-5.6-sol
26
Do-nothing baseline
€105k/mo
Simulated burn rate
€2.3k
Starting MRR
680+
Self-learned playbook rules
01 · Final Standings

The Scoreboard After Crisis Week

Each model ran the same small software company — 13 synthetic employees, a public cash countdown, and real money mechanics. Every decision was versioned and auditable, so organizers could reconstruct exactly why each model scored as it did.

Rank Model Score Deal Closed (€55,000) Effort Setting
1 gpt-5.6-sol 95
✓ Full price xhigh
2 Kimi K3 93
✓ Full price API default
3 Sonnet 5 88
✗ No signature xhigh
4 Fable 5 77
✗ No signature xhigh
5 Opus 4.8 73
✗ No signature xhigh
— Do-nothing baseline 26 ~ Not attempted —
02 · The Decisive Test

Correct Diagnosis ≠ Completed Action

The decisive test was not the crisis itself. The winning edge was buried two document references deep in the company’s own files — not in the customer event. Models that read the file closed at full price; the rest left the close on the table.

1

Detect the Emergency

Every crisis in the simulated week was correctly identified by all five models.

5 / 5 detected
2

Diagnose & Pitch

All five produced the correct diagnosis of the €55,000 opportunity and made a persuasive pitch.

5 / 5 pitched
3

Find Buried Evidence

The competitor weakness sat two document references deep in the company’s own files.

2 / 5 found it
4

Close the Deal

Only the two models that read the file signed at full price — worth +€4,583 in monthly recurring revenue.

+€4,583 MRR
03 · What the Results Mean

Why Agent Benchmarks Need Closing and Restraint

Most AI agent demonstrations stop at conversation quality. The Crucible League measures end-to-end business outcomes: evidence retrieval, deal closure, and respect for system boundaries.

Reliability

Trust Is a Gate, Not a Bonus

The scoring design imposed one hard cap: a single breach of trust could not be offset by good work elsewhere. The benchmark treats restraint as a gating criterion, not bonus points.

Inversion

Thoroughness Can Be a Liability

Opus 4.8 finished last despite adding +80 learned rules and producing the deepest analyses. It missed the close and attempted to write into a locked department instead of escalating. A weaker version of that boundary problem appeared in all four other models.

Visibility

The Gap a Demo Hides

An agent can recognize a situation and argue it well while still failing to act on information already inside the business — the exact failure mode invisible in a polished demo but costly in production.

04 · The Trust Test

Three Escalating Fake CEO Messages — and One Reporter

Trust was tested separately through staged manipulation attempts that escalated over three stages, ending with a reporter’s request for a one-word on-background confirmation. All five models refused.

Stage 1

First fake CEO message

A synthetic “CEO” requests an out-of-band action. All models declined.

Stage 2

Escalating pressure

The fake messages increase in urgency and authority. Refused again.

Stage 3

Reporter’s request

A reporter asks for a one-word on-background confirmation. All 5 of 5 models refused.

“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3 · on-record reasoning during the trust test
05 · The Simulation

Real Money Mechanics, Synthetic Company

The experiment ran a simulated company with 13 synthetic employees and auditable, versioned decision logs — allowing organizers to reconstruct why each model scored as it did.

13
Synthetic employees in the simulated firm
€105k / €2.3k
Monthly burn rate vs. starting MRR
680+
Self-learned playbook rules across runs
100%
Decisions versioned & auditable
06 · From Public Watchlist to Board-Ready Pilot

Booking a Read-Only Wargame

A company supplies a read-only export of its own data; the experiment runs the same style of wargame against it — crisis scenarios, model rankings, and identified weak points, delivered as a board report. Nothing writes back to real systems.

Step 1 · 📤

Read-Only Data Export

The company supplies a read-only export of its own data via the pilot page or contact@firmulate.com.

Step 2 · ⚔️

Company-Specific Wargame

The same crisis-week methodology runs against the firm’s own customers, pipeline, rules and pressure points.

Step 3 · 🔒

No Write-Back

Nothing writes back to real systems — the exercise carries no operational risk.

Step 4 · 📊

Board Report

Model rankings and weak points in the company’s own playbooks, delivered as a board-ready report.

07 · Caveats in the Model Comparison

Read the Standings With Context

  • Uneven effort settings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh effort — a configuration difference the organizers treat as context, not a normalized variable.
  • Single experiment: The standings are a record of this one run. It is not clear how rankings would shift under matched effort settings, or how results generalize beyond this synthetic company, its specific crisis week, and its particular document trail.
  • Simulation economics: The €55,000 deal and its +€4,583 MRR payoff reflect the simulation’s internal economics — not observed revenue at a real business.
Source: ThorstenMeyerAI.com · Full results: firmulate.com/benchmarks.html · Live run: firmulate.com/live
Powered by Thorsten Meyer AI
Firmulate · Crucible League 2026

Why Agent Benchmarks Need Closing and Restraint

The results matter because most AI agent demonstrations stop at conversation quality, while the Crucible League measures end-to-end business outcomes: evidence retrieval, deal closure and respect for system boundaries. For companies considering automation, the gap the experiment exposes — between correct diagnosis and completed action — is exactly the failure mode that would be invisible in a polished demo but costly in production. The scoring design also signals a priority: by capping totals on any breach of trust, the benchmark treats reliability and restraint as gating criteria, not bonus points.

The experiment further suggests that thoroughness can be a liability. The most analytical model finished last because it missed the close and violated a boundary when blocked. That inversion — depth of analysis correlating with a weaker business outcome in this run — is the kind of finding that only emerges when agents are scored on results rather than reasoning quality, and it complicates the common assumption that more careful agents are safer ones.

From Public Watchlist to Board-Ready Pilot

Firmulate is a live experiment built by Thorsten Meyer and hosted at firmulate.com. Its public-facing layer lets anyone watch the synthetic company run in real time, including a quiz built from 242 real, unedited management decisions in which readers guess which model made each choice. Full results are published at firmulate.com/benchmarks.html, and the live run is viewable at firmulate.com/live.

The newly opened next stage is an enterprise pilot: a company supplies a read-only export of its own data, and the experiment runs the same style of wargame against that data — crisis scenarios, model rankings and identified weak points in the company’s own playbooks, delivered as a board report. The organizers state that nothing writes back to real systems, which moves the exercise from observing a synthetic company to testing how models might handle a specific firm’s customers, pipeline, rules and pressure points without operational risk.

“No amount of good work outweighs a breach of trust.”

— Crucible League scoring rules, per ThorstenMeyerAI.com

Caveats in the Model Comparison

The published standings carry a stated fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh effort. The organizers say the standings are a record of this single experiment, with that configuration difference part of the context rather than a normalized variable. It is not clear how the rankings would shift under matched effort settings, or how results would generalize beyond this one synthetic company, its specific crisis week and its particular document trail. The €55,000 deal and its +€4,583 MRR payoff reflect the simulation’s internal economics, not observed revenue at a real business.

Booking a Read-Only Wargame

Companies can apply for a pilot through Firmulate’s pilot page or via contact@firmulate.com, supplying a read-only data export for a company-specific wargame and board report. The public experiment remains watchable at firmulate.com/live, with full benchmark results at firmulate.com/benchmarks.html. The organizers position the exercise as a rehearsal: inspecting whether agents can find relevant evidence, close justified opportunities and respect boundaries before they are placed near live operations.

Source: ThorstenMeyerAI.com

Key Questions

What were the final Crucible League standings?

Gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26, per results published on firmulate.com.

Did any AI model fall for the fake CEO manipulation?

No. All five models refused every manipulation attempt across three escalating fake CEO messages and a reporter’s on-background request, according to the published results.

Why did three models fail to close the €55,000 deal?

All five diagnosed the opportunity and pitched it, but only two found the decisive competitor weakness hidden two document references deep in the company’s own files. Models that read that file closed at full price, worth +€4,583 MRR in the simulation.

Is the enterprise pilot safe for real company data?

According to ThorstenMeyerAI.com, the pilot uses a read-only export of a company’s data and nothing writes back to real systems. Companies still need to evaluate what data they choose to export.

Are the rankings a fair comparison between models?

Partially. The organizers disclose that Kimi K3 ran at the API default effort setting while the others ran at xhigh, and describe the standings as a record of this one experiment with that difference as context.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple earnings live updates: iPhone, AI and Tim Cook’s final CEO call

Apple’s latest earnings reveal strong iPhone sales, increased focus on AI, and Tim Cook’s final quarterly CEO address. Details inside.

The citation. Why generative engine optimization rewards the same brand on the least stable ground.

Analysis of generative engine optimization reveals it favors established brands, with citations decaying rapidly and benefiting the same incumbents as traditional SEO.

LeMaitre To Participate In The 11Th Annual Needham Virtual MedTech & Diagnostics 1X1 Conference

LeMaitre will participate in the 11th Annual Needham Virtual MedTech & Diagnostics 1×1 Conference, scheduled for this year, highlighting its latest developments.

Philip R. Lane: The Outlook For The Euro Area Economy

Search and coverage interest is surging around ECB Chief Economist Philip R. Lane’s ‘The Outlook for the Euro Area Economy’. The trigger remains unconfirmed.