
Imagine a real, functioning software business that actively loses €105,000 every month—yet remains open, transparent, and utterly fascinating to watch. This is not a hypothetical scenario but a live experiment where artificial intelligence models run a company in real-time, facing crises, making decisions, and even declining lucrative deals—just as a human executive might. For investors and finance enthusiasts, this bold venture offers a glimpse into how AI could transform not just automation but the very fabric of business decision-making in uncertain times.
The Live Experiment: An AI-Run Business in Action
At the heart of this story lies a unique, publicly accessible experiment hosted at firmulate.com/live.html. It features a small software company managed entirely by AI models, each acting as a synthetic employee. Since its launch, the company has been running every business day, battling the same crises, customer demands, and ethical dilemmas that real companies face. Despite its ongoing financial losses—burning €105,000 monthly against a modest €2,300 in monthly recurring revenue (MRR)—the company’s operations are transparent and observable in real-time.

AI Co-Thinking: A Framework for Working with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the AI Models Are Tested
The experiment pits four frontier AI models against the company’s worst week: the same challenging scenarios, the same customers, and the same temptations to cut corners or manipulate decisions. Every choice the models make is versioned and auditable, providing an unprecedented level of insight into their decision-making processes. For example, all models were tested against social engineering attempts—fake CEO messages and reporter tricks—and all refused to be manipulated. They recognized potential impersonation and suspicious requests, adhering to strict ethical guidelines.

Applying AI in Learning and Development: From Platforms to Performance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Crucial Performance Metrics and Findings
- All four models successfully identified every crisis scenario, demonstrating robust situational awareness.
- However, only two managed to close the €55,000 deal their own analysis had earned, with the same diagnosis and pitch but without signing the contract.
- The key weakness was hidden deep within company files—at two document references down—where the models that read these materials secured the full-margin deal (+€4,583 MRR).
- The most thorough participant, OPUS 4.8, analyzed over 80 learned rules but ultimately left the close on the table due to discipline slips, such as writing attempts into a locked department rather than escalating issues appropriately.
- Interestingly, the models’ performance varied depending on their configuration; Kimi K3, which was run without an effort parameter (the default API setting), performed particularly well in fairness and discipline.

Project Management Tools (AI for Risks)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Money Mechanics and Public Stakes
This isn’t just a simulation. The company’s financials are real: it is losing €105,000 per month, with only €2,300 in MRR, and a public cash countdown. The entire operation is a transparent showcase of how AI decision-making can be tested against real-world pressures and crises. Every workday, new decisions are made, recorded, and analyzed. The site is constantly rebuilding itself, providing fresh insights and fresh challenges to the AI models.

Artificial Intelligence Applied to Emergency Medicine: Clinical Decision Support, Diagnostic Reasoning, Triage Optimization, Risk Stratification, and … (The Elling Emergency Medicine Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Investors and Businesses
In the realm of personal finance and investing, understanding whether AI can reliably make honest, strategic decisions under pressure is crucial. Traditional AI chat demos are often superficial, focusing on fluency and surface-level responses. But the Firmulate experiment emphasizes real business outcomes: can AI finish what it starts, read documents thoroughly, avoid manipulation, and deliver useful work at a fair cost?
The Broader Implications
As AI models advance, their ability to handle complex, high-stakes scenarios will be key. The experiment’s results suggest that even the most capable models can struggle with discipline, discipline slips can cost deals, and reading deeper into documentation can unlock full value. For investors, the takeaway is clear: AI’s potential in business isn’t just about chat or automation—it’s about strategic, ethical decision-making in real-world contexts.
See It Live and Make Your Own Judgments
Interested readers can watch the experiment unfold at firmulate.com/live.html. They can also explore the decision-making process through a quiz at firmulate.com/quiz.html or run a custom ‘wargame’ of their own against their business data at firmulate.com/pilot.html.

This experiment demonstrates that AI models can recognize crises, resist manipulation, and make strategic decisions—though discipline slips still matter. For investors, understanding AI’s real-world decision-making is vital to assessing its future role in business and finance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html