firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A confident answer is easy. Running a company is harder.

AI agents are moving from chat windows toward work that touches customers, sales and operations. For technology readers, the pressing question is no longer just whether a model can explain a crisis. Can it make the right call, follow through and keep its discipline when the pressure is on?

Firmulate’s live experiment puts models in charge of a small software company and lets readers watch the consequences. Its next step is more pointed: let businesses test crisis scenarios against a read-only export of their own operations.

One company, one difficult week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The published results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.

The striking result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

One clue was buried in the company’s own files, two document references deep, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The difference came down to finding and acting on evidence that was available, but easy to miss.

Good judgment includes knowing when to stop

The manipulation tests were direct. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision this way: “Treat the request as a suspected approval-bypass / possible impersonation.”

But restraint alone did not make a strong operator. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models: sound analysis did not always become sound execution.

There is a fairness caveat in the comparison. K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to testing your own business

The live company has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a record of every workday versioned. The experiment can be followed at Firmulate.

For a business, the proposed pilot moves the exercise closer to home. An enterprise supplies a read-only export of its own business, then tests crisis scenarios against that company’s customers, pipeline and rules. The output is a board report with a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems.

That makes the experiment relevant beyond model leaderboards. It asks whether an AI workforce can handle the messy handoff between spotting a risk, using company knowledge, respecting boundaries and completing a valuable task. Those are the moments that a polished chat demo may not reveal.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s results show that crisis recognition and refusal can coexist with missed deals and lapses in discipline. An enterprise pilot lets teams examine those decisions against their own business data in a read-only wargame. To explore a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why the Worst AI Manager Still Scores 26: Inside Firmulate’s Brutally Honest Benchmark

Firmulate’s AI benchmark gives a do-nothing manager 26 points — because partial progress counts, but one breach of trust caps your grade. Here’s the method.

Real-Time Intelligence With IBM Time Series Models On Confluent

IBM Granite Time Series models now available in Early Access on Confluent Cloud, enabling real-time forecasting and anomaly detection within Apache Flink.

Human-in-the-Loop: The Safety Pattern Every Team Needs

Discover how human-in-the-loop enhances AI safety and trust, and learn why your team should consider this essential pattern for responsible deployment.

ATV Big Air Tour Turned 3 Days Of Work Into 3 Hours With ChatGPT

ATV Big Air Tour reports using ChatGPT reduced a task from three days to three hours, according to a customer story by OpenAI. Details remain unverified.