
For anyone weighing a financial decision, trust matters as much as a confident recommendation. That is also the question facing businesses considering AI: will an automated workforce handle pressure responsibly, and follow through when the stakes rise? Firmulate’s live experiment puts AI models in charge of the same small software company through its worst week, with real money mechanics and crises. The results offer a practical case for testing how AI behaves before giving it a role in business decisions.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
The experiment gave each frontier model the same customers, crises and temptations. Every decision was versioned and auditable. The company is synthetic, but its financial mechanics are real: 13 synthetic employees, a monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown. Its playbooks include more than 680 self-learned rules, and every workday is versioned. The live company is watchable at Firmulate.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The experiment’s standard is deliberately sharp: partial progress counts, but one breach of trust caps the total. As the rules put it, “no amount of good work outweighs a breach of trust.”
Seeing the crisis was not enough
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. They made the same diagnosis and delivered the same pitch; some still left without a signature. That gap between understanding a decision and acting on it is difficult to spot in an ordinary chat demonstration.
The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result points to a familiar business challenge: useful information may be available, but an AI still has to find it and carry its analysis through to a sound action.
Integrity was tested directly, too. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as “a suspected approval-bypass / possible impersonation.” The episode shows why organizations need to assess how an AI responds to pressure to bypass safeguards, not just whether it can complete routine tasks.
Strengths can come with blind spots
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, while discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. Thoroughness alone, the experiment suggests, does not guarantee sound judgment or reliable execution.
There is a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The published ranking therefore comes with a meaningful difference in how the models were run.
For readers thinking about AI in finance or business, the useful question is not simply whether a model sounds persuasive. It is whether it can locate relevant information, respect boundaries, resist manipulation and complete a decision when the evidence supports it. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com.
From watching to a company-specific test
The next step is to test models against an organization’s own circumstances. Firmulate’s proposed enterprise pilot uses a read-only export of a company’s business to stage crisis scenarios and produce a board report with model rankings and weaknesses in the company’s playbooks. Nothing writes back to real systems. That gives decision-makers a way to examine how an AI handles their own information and pressures before considering where it should be used.

Put your own playbooks to the test
Watching a live company reveals how AI models behave under pressure. A pilot can make that test relevant to your own business by using a read-only export and reporting where models succeed or where your playbooks have weak points. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
