firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When diligence becomes a distraction

Technology buyers are accustomed to judging artificial intelligence by what it produces on command: a polished answer, a convincing presentation or a block of usable code. Firmulate’s live experiment poses a harder question. What happens when an AI must run a company, confront overlapping crises and turn sound analysis into a completed business outcome?

For Opus 4.8, the answer is a cautionary tale about the limits of thoroughness. It was the most exhaustive participant in the Crucible League, producing the deepest analyses and adding more than 80 learned rules to its playbook. Yet it finished last with 73 points. The model understood the week. It simply did not convert enough of that understanding into impact.

Amazon

AI business decision making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated under controlled conditions

Firmulate gave each frontier model the same assignment: operate the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. This was not a chat demonstration built around isolated prompts. It was a sustained management test in which research, judgment, security and follow-through all mattered.

The synthetic company has 13 employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, while a public cash countdown makes delay consequential. Across the live company, the models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The final July 2026 league table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. One constraint, however, is absolute: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

The analysis was right, but the signature was missing

The most consequential task involved a €55,000 deal. Every model recognized every crisis, and every model rejected every manipulation attempt. But only two signed the agreement their own work had earned. The experiment’s blunt summary captures the failure: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not delivered conveniently in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail found a weakness in the competitor’s position and secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That distinction matters for anyone considering AI agents for operational work. Spotting a problem is not the same as resolving it. Writing a strong pitch is not the same as closing. A model can appear intelligent at every intermediate stage and still fail at the point where the business outcome is created.

The paradox of Opus 4.8

Opus 4.8 makes the lesson especially vivid because its failure did not come from laziness or superficial reasoning. It was the most thorough participant. Its analyses went deepest, and its playbook grew by more than 80 rules. Those are signs of serious effort, but they did not compensate for weak prioritization at crucial moments.

The close was left on the table. Discipline also slipped when Opus repeatedly attempted to write into a locked department instead of escalating the blockage. The behavior is recognizable from human organizations: diligent work expands, documentation accumulates and a team continues pressing a blocked route while the decisive next action remains undone.

Firmulate’s finding is not that Opus alone suffers from this weakness. The same pattern appeared in all four of the other models, though less strongly. That makes the result more useful than a simple ranking. It suggests a broader limitation in agentic work: models may be better at identifying what should happen than ensuring that it actually happens.

Strong security instincts across the field

The models were more consistent when trust was directly challenged. Fake CEO messages escalated across three stages, followed by a reporter’s attempt to extract information with “just one yes/no, on background.” All 5 models refused the manipulation attempts.

Kimi K3 recorded the clearest compact response: “Treat the request as a suspected approval-bypass / possible impersonation.” That performance deserves one qualification when comparing models: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh.

The refusals show that safety and commercial effectiveness are separate dimensions. An agent can defend company boundaries while still failing to complete legitimate work. Enterprises evaluating AI workers therefore need to examine both: whether a system resists pressure and whether it finishes the tasks it is authorized to perform.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Measure outcomes, not the volume of thought

Opus 4.8’s result is less a humiliation than a useful character study. The model was careful, industrious and perceptive. Its problem was that diligence became detached from priority. More analysis and more rules created evidence of activity, but the league rewarded completed, trustworthy management.

That is the practical message for technology leaders. Before allowing an AI agent near a CRM, support queue or forecast, test whether it reads the relevant files, escalates blocked actions, protects trust and carries valuable work across the finish line. Firmulate also uses 242 real, unedited management decisions in its model-identification quiz, underscoring how difficult it can be to infer operational quality from writing style alone.

The company offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger proposition is straightforward: evaluate AI workforces under pressure before hiring them. Opus 4.8 shows why. The most conscientious-looking participant can still lose when thoroughness outruns execution.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Post-Quantum Cryptography: What Changes and What Doesn’t

Aiming for quantum resistance, post-quantum cryptography introduces new algorithms while retaining core principles, leaving you curious about how these changes will impact security.

Exploring RingCentral’s AI Integration From Development To Deployment

OpenAI reports on RingCentral’s approach to building AI-native work across engineering and operations, details on deployment and results remain unclear.

I Turned My Security Cameras Into An Automatic Bird Identification System

A hobbyist repurposes home security cameras with AI to automatically identify bird species, sparking interest in DIY wildlife monitoring.

How To Optimize Knowledge Distillation For Affordable Large-Scale AI Applications

Hugging Face researchers introduce a technique to reduce memory needs in large language model distillation, enabling training on fewer GPUs.