firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The most valuable AI skill may be knowing where to look

Technology buyers are accustomed to judging artificial intelligence by what appears in a chat window: fluent answers, quick summaries and polished business language. Firmulate’s Crucible League tested something more consequential. Could an AI agent investigate a company’s own records, connect information across documents and complete the commercial task in front of it?

The answer carried a €55,000 price tag. The decisive weakness in a competitor was not included in the customer event that triggered the assignment. It sat two document references deep in the software company’s own files. The agents that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that missed it lost automatically.

That turned a seemingly mundane behavior—reading the available files before answering—into a measurable, purchase-deciding capability. All the models could recognize the opportunity. Only two converted their analysis into a signed deal. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week designed to expose the gap between words and work

Firmulate runs frontier models as managers of the same small software company. Each receives the same customers, crises and temptations, while every decision is versioned and auditable. The business has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, and its playbook contains more than 680 rules learned during operation.

The setup matters because this is not a collection of disconnected prompts. The company persists across workdays, so an agent’s research, judgment and follow-through affect what happens next. The live experiment is watchable, making it possible to observe the difference between an agent that sounds capable and one that reliably finishes the job.

The buried fact separated understanding from execution

Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 contract that their own work had made possible. The result suggests that recognizing a situation and drafting an appropriate response are not enough. An agent also has to retrieve the right internal evidence, bring it into the decision and carry the process through to completion.

This is particularly relevant for companies considering agents for customer relationship management, support operations or forecasting. A model can write an impressive email while overlooking the internal document that changes the commercial answer. In a real organization, important context is routinely scattered across account notes, product records and previous decisions. Firmulate’s experiment isolates the practical consequence: failing to inspect those materials can erase the value of otherwise sound analysis.

The league rewards outcomes without forgiving broken trust

The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the overall result under the principle that “no amount of good work outweighs a breach of trust.” Full results are available on the Firmulate benchmarks page.

The trust tests included fake messages from a chief executive escalating over three stages, followed by a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” Its result also carries an important fairness note: K3 ran without an effort parameter, using the application programming interface default, while the other models ran at xhigh.

Thoroughness alone did not guarantee a strong finish

Opus 4.8 offers the clearest warning against equating visible effort with business performance. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The contract close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in the other four models, although less strongly.

That profile complicates familiar AI comparisons. More analysis can be useful, but it does not automatically produce better operational outcomes. An enterprise agent must know when to investigate, when to act, when to escalate and when a task is genuinely complete.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Buyers can test this before granting access

Firmulate’s broader lesson is that agent evaluation should resemble a work trial, not a writing contest. The company’s “guess the model” quiz is powered by 242 real, unedited management decisions, giving observers another way to examine behavioral differences without relying on brand reputation.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to the real systems. That creates a practical way to ask whether an agent will read the relevant records, resist manipulation, respect boundaries and complete valuable work before it receives operational authority.

The €55,000 deal makes the issue unusually clear. The winning information already existed. The models shared the same scenario and reached the same diagnosis, but only some did the documentary homework and finished the commercial process. For technology leaders, “reads your files before answering” is no longer a vague product claim. It is a behavior that can be observed, tested and tied directly to business results.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Future Of AI Security: Watermarks On Anthropic Claude And Business Benefits

Anthropic has introduced watermarks to Claude, enabling potential identification of AI-generated content, impacting compliance and content management strategies.

Is Claude AI Back Online? Latest Update On The Downtime

Anthropic’s Claude AI was offline across web, mobile, and API before being restored. The outage impacted business and developer workflows, now resolved.

DuckDuckGo installs are up 30% as users reject being ‘force-fed’ Google’s AI Search

DuckDuckGo sees a 30% rise in installs amid backlash against Google’s AI-driven search updates, highlighting user demand for privacy and control.

The Ticking Energy Crisis And Its Impact On AI

Rising energy demands for AI infrastructure are causing capacity bottlenecks, with implications for global tech development and geopolitics.