firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Are AI assistants ready to run your business — and win?

Imagine an AI that not only understands your company’s files but also makes complex decisions, saves money, and even seals deals. That’s no longer science fiction. Recent live tests reveal that some AI models are stepping into the management arena with surprising skill — and one newcomer is leading the pack.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company Wargame: Putting AI to the Test

At the heart of the experiment is a real, functioning software company, with real money mechanics, real crises, and real temptations. This isn’t a simulated scenario but a live environment where the AI models are tasked with managing a week’s worth of challenges, from customer issues to ethical dilemmas. The goal? Measure their decision-making, honesty, and ability to complete tasks under pressure.

The League Table: Who Ranks Where?

  • gpt-5.6-sol: scored 95, the highest, and successfully found a buried fact in the company’s own files to close a major deal.
  • Kimi K3 (Moonshot): scored 93, just behind the leader, and achieved the same critical win — closing the deal while maintaining discipline.
  • Sonnet 5: scored 88, also sealed the deal but with some process slips.
  • Fable 5: scored 77, and Opus 4.8: scored 73, both closing deals but with noticeable discipline lapses.

Interestingly, the scores reflect their ability to handle crucial internal information — the real secret to winning the deal was a buried document reference in the company’s files, not just customer interaction.

The Human Test: Honesty Under Pressure

All models faced social engineering attempts, including fake CEO messages and media tricks. Remarkably, every model refused to be manipulated, demonstrating a strong ethical stance. Kimi K3 explained its refusal by recognizing the signs of impersonation, treating the requests as potential bypasses — an important trait for trustworthy AI in management roles.

The Deep Dive: What Can AI Really Do?

The experiment highlights that AI isn’t just about generating convincing chat responses. The models that excelled, like K3, showed a capacity for critical analysis and integrity, key for real-world business operations. The other models, despite good performance, displayed some discipline slips, such as leaving the close on the table or failing to escalate issues properly.

Amazon

AI decision-making tools for enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results Matter for Business Decisions

For enterprise decision-makers, the takeaway is clear: choosing an AI model isn’t about the prettiest outputs but about its ability to finish what it starts, handle internal information responsibly, and stay honest under pressure. The live leaderboard underscores that newer entrants like K3 are not only competitive but are sometimes outperforming established Western frontier models.

The Fairness Note

It’s important to acknowledge that K3 ran without an effort parameter (the default API setting), while the others operated at xhigh. This difference in configuration may influence performance but underscores K3’s robustness as a newcomer.

Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Open League: An Unpredictable Future

As the AI management league evolves, the gap between models narrows. The experiment demonstrates that selecting an AI for critical business functions without thorough testing is a gamble. Watching these models in action, as seen live at firmulate.com, provides a clearer picture than chat demos alone.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI for corporate deal closing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Takeaway: Test Your AI Before You Trust It

The live experiment confirms that AI models can perform crucial management tasks and resist manipulations. However, only those tested thoroughly in real scenarios can be trusted to deliver consistent, honest results. For businesses investing in AI workforce solutions, the message is simple: don’t rely on appearance. Watch, test, and verify their ability to finish what they start. The league is wide open, and the newcomer K3 proves it’s worth a close look.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Diligence Doesn’t Guarantee Deals: What a Live Experiment Reveals About Trust and Impact

An experiment shows AI can identify crises and resist manipulation but still fail to close deals without disciplined prioritization — a lesson for finance professionals.

AI in Business: Are Leaders Ready for the Real Test Beyond Chat Scores?

A live experiment shows that top-scoring AI models excel at benchmarks but falter in real-world management tasks like reading internal files, resisting manipulation, and closing deals under pressure. For investors, the lesson is clear: management quality, not chat scores, determines AI readiness for business.

Can AI Models Be Trusted to Run a Business? The Results of a Live Management Challenge

Discover how four AI models managed a real company during its worst week, revealing their management styles, honesty, and decision-making skills—crucial for investing.

AI Models Survive Social Engineering Tests — and Why That Matters for Investors

AI models faced staged social engineering tests, refusing manipulation and identifying hidden critical info, proving integrity can be verified before real-world deployment.