
Can you recognize an AI by the decisions it makes?
Technology fans are accustomed to identifying gadgets from camera bumps, screen bezels or benchmark charts. Firmulate poses a stranger challenge: can you identify a frontier AI model from the way it manages a company in crisis?
Its interactive guess-the-model quiz draws from 242 real, unedited management decisions. Readers see how an AI responded to a business situation, choose which model they think was responsible and then discover the answer alongside a character profile. The differences are not merely stylistic. Across an otherwise identical test, the models developed recognizable habits around research, follow-through, discipline and risk.
The result feels playful, but the underlying question is serious. If AI systems are going to operate support queues, customer relationships or forecasts, polished prose matters less than whether they read the available evidence, resist manipulation and complete the work they started.
As an affiliate, we earn on qualifying purchases.
The worst week at the same company
Firmulate placed each frontier model in charge of the same small software business during its worst week. Every participant encountered the same customers, crises and temptations. Every workday and decision was versioned and auditable, allowing differences in behavior to be compared rather than explained away by changing circumstances.
The simulated company has 13 synthetic employees and deliberately unforgiving economics: it burns €105,000 per month against €2,300 in monthly recurring revenue. A public cash countdown keeps the pressure visible, while the business has accumulated more than 680 playbook rules learned through experience.
In the final July 2026 Crucible League table, gpt-5.6-sol finished first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 recorded 73. For comparison, a do-nothing baseline scored 26 because partial progress still counted.
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
- Do-nothing baseline: 26
The evaluation also imposed a hard ethical boundary: a single breach of trust capped the total because, as the experiment states, “no amount of good work outweighs a breach of trust.”
Everyone saw the trouble; not everyone finished
All the models detected every crisis, and every one resisted every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That finding separates understanding from execution. A model can identify a commercial opportunity, assemble a convincing case and still fail at the final managerial act. In a conventional chat demonstration, the analysis might look excellent. In a running company, the missing signature is the outcome.
The decisive evidence was easy to overlook. A competitor weakness was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that followed the references and read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This makes the quiz more revealing than a simple writing-style game. The recognizable traits emerge from what each model notices and completes: whether it searches beyond the immediate event, turns knowledge into action and follows operating discipline when the week becomes noisy.
Different management personalities emerge
Opus 4.8 provides the clearest cautionary profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last among the competing models. The close was left on the table, and it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the other participants, though less strongly.
Kimi K3 displayed a different kind of clarity during the social-engineering test. The company received fake CEO messages that escalated across three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 models refused. K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”
There is an important fairness qualification when comparing results. K3 ran using the API default because it had no effort parameter, while the other models operated at xhigh. Its score therefore belongs in the table, but readers should understand that the test settings were not identical on that dimension.

A benchmark you can interrogate
Firmulate’s experiment turns model evaluation into something readers can inspect rather than simply accept. The live company is real software with real money mechanics, and its continuing decisions are watchable. The quiz makes that evidence approachable by asking visitors to develop their own sense of each model before revealing the identity.
The broader lesson is that AI management ability does not collapse into a single talent. Thorough analysis can coexist with weak follow-through. Strong crisis detection does not guarantee a completed sale. Shared resistance to social engineering can conceal meaningful differences in research and operational discipline.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to their real systems, allowing organizations to test an AI workforce against familiar conditions before giving it operational responsibility. For technology buyers, that may be the most useful twist: the model with the most impressive answer is not necessarily the manager that gets the job done.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html