
Just like in fitness, where strength and discipline determine your progress, the true test of AI in business isn’t how well it chats—it’s how reliably it manages real-world crises under pressure. A recent live experiment with AI models running a simulated company reveals that scoring well on benchmarks doesn’t necessarily mean the AI can handle the messy, unpredictable realities of management.
How AI Performs Under Pressure: The Firmulate Experiment
Imagine a small software company battling a week of crises—customer churn, pricing shocks, PR nightmares. Now, picture AI agents stepping into the shoes of managers, making decisions to save the day. That’s exactly what the team at Firmulate did with their live benchmark, pitting four state-of-the-art AI models against the same challenging scenario.
All four models—ranging from GPT 5.6 to emerging newcomers like Kimi K3—were tasked with navigating the company’s worst week. They faced the same customers, the same crises, and the temptation to cut corners or manipulate systems. Each decision was recorded and made auditable, mimicking real-world management complexity.
Results: The Scorecard of AI Management
- GPT-5.6: Scored 95 out of 100, identified critical information buried two documents deep, and successfully closed a €55,000 deal—covering the full scope of what management requires.
- Kimi K3: Scored 93, also closed the deal, and demonstrated the cleanest discipline among the models. Notably, it refused all manipulation attempts, including fake CEO messages and reporter tricks.
- Sonnet 5: Achieved 88, closed the deal but with some process slips, like leaving the final offer on the table and slipping discipline under pressure.
- Opus 4.8: Scored 73, closed the deal but showed weaknesses similar to Sonnet, with less thorough analysis and some slip-ups in escalation processes.
The key takeaway? While all models identified crises and refused manipulation, only two—the top two—finalized the deal based on their own analysis. The others hesitated or left money on the table, revealing a crucial gap: true management quality is about more than just surface-level chat responses.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Beyond the Surface
The real edge came from reading deeper into the company’s own files, not just reacting to customer events. Models that could access and interpret critical internal documents won the full price deal, adding over €4,500 monthly recurring revenue (MRR). This highlights a vital aspect of management: the ability to leverage internal data, not just respond to external crises.
Social Engineering Resistance
The models faced a staged social engineering attack—fake CEO messages escalating in three stages, plus a reporter trick requesting a simple background quote. All five models refused to be manipulated, citing concerns about impersonation or bypassing approval protocols. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
internal data analysis software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: Money, Discipline, and Human-Like Judgment
Firmulate’s experiment isn’t just virtual. The live company, with 13 synthetic employees, runs actual money mechanics—burning €105k per month against just €2.3k in MRR. It’s a real testbed where every decision impacts the bottom line, and every day the AI models are versioned, analyzed, and challenged.
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, still left opportunities unexploited and discipline slipping—leaving money on the table. This underscores that even the most advanced models struggle with consistent, disciplined management under stress. Performance isn’t just about scoring well in demos; it’s about staying honest, reading deeply, and executing with discipline.
social engineering resistance training for managers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway for Business Leaders
For companies eyeing AI to handle support, CRM, or forecasting, the question isn’t just about how well it writes or responds. It’s whether the AI can finish what it starts, read and interpret internal data, and remain honest and disciplined when under pressure. Benchmarks that focus solely on chat quality miss this vital dimension of management capability.
As the leaderboard from Firmulate shows, models like GPT 5.6 and Kimi K3 excel not only because they diagnose crises but because they act decisively and ethically, even when tested. Conversely, even the most thorough models can stumble if they’re not disciplined enough or don’t read deeply enough into internal files.
As an affiliate, we earn on qualifying purchases.
Beyond the Test: Wargaming Your AI Workforce
Curious how your own AI systems would perform in real management scenarios? Firmulate offers a live, real-time platform where enterprises can run their own scenarios—an AI wargame, with no impact on actual systems. It’s a chance to see whether your AI team can truly handle the complexities of real business crises before you hire or deploy them at scale.

The real measure of AI management isn’t just how well it chats. It’s whether it can read deeply, stay disciplined under pressure, and finish what it starts—crucial qualities that benchmarks often overlook but are vital for real-world success.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html