
Imagine having an AI team managing your business during its worst week—making critical decisions under pressure, without a single slip. Would you trust these artificial managers? As AI tools become increasingly integrated into finance and investing, understanding their management personalities isn’t just interesting—it’s essential. Today, we explore a groundbreaking experiment where four advanced AI models ran a real, money-losing software company through its most turbulent week, revealing surprising insights into their decision-making styles and trustworthiness.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI to the Test
In a first-of-its-kind live demonstration, four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with managing a real company facing a series of crises. Unlike typical chat demos, this was a full simulation of a small software business with actual money mechanics, a team of synthetic employees, and a public cash countdown. The models operated in real-time, making decisions across a week of simulated but plausible events, including customer crises, internal conflicts, and manipulative social engineering attempts.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: What These AI Managers Revealed
Despite their different personalities and approaches, all four models successfully identified every crisis and refused manipulation attempts. For example, when fake CEO messages escalated through multiple stages, all models refused to act on them, citing concerns about impersonation and bypassing approval processes. This shows a consistent core of honesty and integrity across the models.
However, differences emerged in their ability to close deals and follow through on decisions. Only two models—gpt-5.6-sol and Kimi K3—signed a €55,000 deal that their own analyses had earned, after assessing the situation thoroughly. The other two, Sonnet 5 and Fable 5, left the deal on the table. Interestingly, the decisive advantage lay not in the initial crisis detection but in how deep each model read into the company’s own files. The models that examined information buried two document references deep were able to identify a hidden, critical fact and secure the deal at full price, worth an additional €4,583 monthly recurring revenue.
As an affiliate, we earn on qualifying purchases.
Personality Profiles in AI: Different Styles, Same Integrity
Looking at the models’ management personalities, some stood out more than others. For instance, Opus 4.8 was the most thorough participant, analyzing over 80 learned rules and conducting the deepest assessments. Yet, it finished last, leaving the close on the table and slipping into a more siloed approach by writing attempts into a locked department rather than escalating. Meanwhile, Kimi K3 ran without an effort parameter, making it more disciplined and fair, which contributed to its successful deal closure. The differences highlight how model configurations influence behavior—some are more meticulous, while others are more decisive.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Investing
This experiment isn’t just a tech curiosity. It underscores a vital truth for anyone in finance or investing: when AI systems are managing or supporting critical decisions—be it in risk assessment, client management, or operational oversight—their management style, thoroughness, and honesty matter immensely.
For investors and finance professionals, the takeaway is clear: it’s not just about whether an AI can generate convincing chat or reports. The real question is whether it can finish what it starts, interpret deep internal data, and stay honest under pressure. These traits directly impact the ROI of AI investments and the trustworthiness of automated decision-making systems.
As an affiliate, we earn on qualifying purchases.
Why Trust Matters More Than Ever
The live company’s mechanics also reveal that, in real-world scenarios, AI models must operate in environments where money is at stake—in this case, losing €105,000 per month against €2,300 in revenue. The models’ ability to refuse manipulative social engineering and identify hidden facts is crucial in safeguarding assets and maintaining integrity. As AI becomes more embedded in financial services, understanding their management personalities can help firms choose the right model for the right task.
See It Live and Make Your Guess
Curious to see how these models perform firsthand? You can test your judgment with the same management decisions at firmulate.com/quiz.html. The live experiment and ongoing company simulation are visible at firmulate.com/live. Here, you can watch this AI-driven management unfold in real-time, observe decision logs, and even run your own scenarios against your business data—without risking your actual operations.
Final Thoughts
As AI models evolve, their personalities—ranging from meticulous to terse—will influence how they handle complex, high-stakes situations. This experiment proves that even the most advanced models can demonstrate unwavering integrity, yet their effectiveness also depends on configuration and depth of data access. For investors and business owners, it’s a call to look beyond surface capabilities and assess how these systems manage trust, honesty, and thoroughness under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.