AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Fitness is built around a simple idea: it is better to discover your limits in training than in the middle of a competition. Businesses using AI face a similar test. A polished demo cannot show what an AI workforce will do when customers leave, a crisis hits, or someone tries to bypass the rules. Firmulate puts those decisions under pressure before a company hands over real work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, on repeat

In the final Crucible League, held in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The experiment asks a practical question: can a model manage a business, not just produce convincing answers?

The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s integrity rule is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

Spotting the crisis is not the same as finishing the job

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The report captured the gap this way: “Same diagnosis, same pitch — no signature.” Recognizing a good move and carrying it through are different tests of management.

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes a useful point for any business: the facts that matter may be present but easy to overlook.

The social-engineering test escalated across three fake CEO messages, then added a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was to “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs judgment

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness detail behind the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a record of this experiment, with that difference disclosed.

From watching to testing your own business

The live company makes the experiment watchable. Its 13 synthetic employees operate with real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. A separate quiz uses 242 real, unedited management decisions to invite readers to guess which model made each choice.

For an enterprise, the next step is a pilot using a read-only export of its own business. Teams can put company-specific crisis scenarios against their own data and receive a board report showing model rankings and weak points in their playbooks. Nothing writes back to real systems. The move is from watching a live-company experiment to seeing how models handle your company’s particular pressures.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

AI readiness is a performance question: can a model protect trust, find the detail that changes a decision and follow through when the stakes rise? Firmulate’s league shows why that deserves a real test. To explore a pilot using your company’s read-only data, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Testing Reveals the Hidden Strengths of Business Discipline — and What Fitness Can Learn

AI models managing a business reveal that true performance under pressure and discipline are invisible in demos but critical for success—lessons for gym enthusiasts alike.

How to Use Range of Motion the Right Way in Training

While mastering proper range of motion enhances your training, understanding the key techniques can prevent injuries and maximize results.

The Most Common Deadlift Setup Errors and How to Fix Them

Keen to perfect your deadlift setup? Discover the most common errors and essential fixes to maximize your strength and safety.

Why Bench Press Setup Matters More Than Most Lifters Think

AIThis post was created with the assistance of artificial intelligence (AI).Your bench…