
A brilliant answer is not the same as a well-run company
Technology buyers have learned to scan coding leaderboards and chat arenas for signs of intelligence. Those tests are useful, but they leave a management-sized hole. They do not show whether an AI agent can triage several crises, follow through on commercial work, resist pressure to break trust and live with the consequences of its choices across days.
Firmulate is testing that gap in public. Its live experiment gives frontier models the same small software company and sends each through its worst week: the same customers, crises and temptations. Every decision is versioned and auditable. The result is less like another chatbot contest and more like a management wargame.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The leaderboard changes when execution counts
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But the benchmark also treats trust as a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”
Those scores matter less than the behavior behind them. Every model spotted every crisis. Every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is the problem conventional AI evaluation tends to miss. Producing a persuasive analysis is answer quality. Turning that analysis into an appropriate, completed business action is management quality. An agent can sound decisive while leaving revenue untouched, or appear cautious while failing to escalate a blocked task.
The winning fact was already inside the company
The decisive clue in the deal was not contained in the customer event. It sat two document references deep in the company’s own files. Models that followed that trail found a competitor weakness, used it in the sales process and won the deal at full price. The result was worth +€4,583 MRR.
This is a more realistic test of enterprise AI than asking whether a model can summarize a document placed directly in front of it. Business work is scattered across customer records, internal notes and prior decisions. The valuable fact may not announce itself as relevant. The agent must recognize that it needs more context, locate it and then act on what it learns.
Safety held up better than follow-through
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding complicates the familiar assumption that safety is mainly about refusing harmful prompts. Here, refusal was the field’s common strength. The sharper separation appeared in ordinary management discipline: reading deeply enough, completing the close and responding properly when organizational boundaries blocked progress.
Opus 4.8 makes the point vividly. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal remained unsigned, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models. More analysis did not automatically produce better management.
One comparison also deserves a fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers should keep that difference in mind when interpreting the final ranking, even though the observed decisions remain auditable.
A company-sized curriculum
Firmulate’s broader proposition is that scenario names such as churn wave, price increase, downround and PR crisis can become a new curriculum for evaluating agents. These situations test prioritization, commercial judgment, institutional memory and honesty toward the board. They expose failure modes that polished conversations cannot.
The live company gives those scenarios consequences. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown keeps running, more than 680 self-learned playbook rules accumulate, and every workday is versioned. The experiment is real, watchable and deliberately uncomfortable.
Readers can also confront their own assumptions through a quiz built from 242 real, unedited management decisions. Guessing which model made each choice is a useful reminder that confident prose is often a poor guide to operational reliability. The full league results and plain-language findings are available on the Firmulate benchmark page.

Evaluate the manager, not merely the chat
For companies considering agents for a CRM, support queue or forecast, the central question is no longer whether the model writes well. It is whether the agent reads the relevant files, finishes what it starts, escalates when blocked and remains honest under pressure.
Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That approach points toward a practical procurement category: test management quality before granting operational authority. Coding ability may get an agent through the interview. A week of consequences reveals whether it should get the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html