
Imagine running your pool business with an AI that not only handles your customer inquiries but also makes vital management decisions during your worst week — and it does so honestly and effectively. Curious how AI models perform under pressure? The live experiment from Firmulate offers a revealing look into their decision-making personalities, with real consequences for a small software company facing its most challenging days.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Introducing the Live AI Business Wargame
At firmulate.com, a pioneering experiment challenges the common perception of AI as merely a chatty assistant. Instead, it tests whether these models can run a business—navigate crises, uphold trust, and close deals—while being observed in real time. The experiment involves four frontier AI models, each running the same small software company, enduring the same troublesome week filled with customer issues, temptations to cheat, and social engineering tricks.
What makes this test unique? Every decision these models make is recorded and auditable. They face real-world crises—like customer disputes and internal document leaks—and are tested on critical moments such as whether they read the company’s files before responding or succumb to manipulation attempts.
As an affiliate, we earn on qualifying purchases.
The Models and Their Scores
- GPT-5.6-sol: Scores the highest at 95, successfully uncovering hidden information in documents and sealing a €55,000 deal.
- Kimi K3: Close behind at 93, demonstrating the cleanest discipline, refusing all manipulation attempts and closing the deal.
- Sonnet 5: Achieves a score of 88, also closing the deal but with some process slips.
- Fable 5: Ranks at 77, with a similar outcome but weaker discipline and missed opportunities.
As an affiliate, we earn on qualifying purchases.
What Did the Models Get Right? And Where Did They Falter?
All four models identified every crisis and refused manipulative tactics, including social engineering attempts—fake CEO messages and background approval requests. For instance, Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation,” showcasing a cautious and security-minded approach.
However, the decisive factor was a buried fact in the company’s internal files—information crucial to winning the deal. Models that read the company’s documents, notably GPT-5.6-sol and Kimi K3, were able to leverage this information fully, closing the deal at the full price (+€4,583 MRR). In contrast, the others missed this detail, leaving money on the table.
AI cybersecurity protection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Personality and Management Style in AI
These results highlight distinct management personalities embodied by each model. GPT-5.6-sol is thorough, detail-oriented, and pays attention to hidden internal data—like a meticulous manager who digs deep before making decisions. Kimi K3 runs with a default API effort level, emphasizing fairness and discipline, and refuses shortcuts or manipulations. Sonnet 5, with more process slips, embodies a slightly less disciplined style, while Fable 5’s weaker discipline leaves potential gains unclaimed.
As an affiliate, we earn on qualifying purchases.
The Real-World Impact and Why It Matters
This live experiment isn’t just about a game; it’s a window into how AI can manage real business operations. The company, a functioning software firm, currently burns €105,000 each month against €2,300 MRR, illustrating the high stakes involved. The decisions made during this crisis week could mean the difference between closing lucrative deals or losing trust and revenue.
Moreover, the models’ refusal to fall for social engineering shows promising honesty and security traits—crucial for AI integration into customer support, CRM, or financial decision-making. The key takeaway? It’s not just how well an AI writes or responds in a chat, but whether it can see the full picture, stay honest under pressure, and complete its work reliably.
How Can Business Leaders Use This Insight?
Before deploying AI into critical workflows, companies can run their own ‘wargame’ simulations with tools like those from Firmulate. You can export your business scenario, observe how your AI models perform in a controlled environment, and understand their management style—whether meticulous, disciplined, or prone to slips. These insights help ensure your AI will do more than just talk well; it will deliver consistent, trustworthy results.
To explore how your business can benefit from such testing, visit firmulate.com/quiz.html and see how your existing AI or future models measure up in these critical management traits.

The live experiment proves that AI models can do more than chat—they can manage crises, uphold trust, and close deals. However, their management personality and discipline level ultimately determine success or failure. For organizations considering AI in management roles, testing like this offers invaluable insight into real-world performance before going live.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.