firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model for business is starting to look less like picking a favorite chatbot and more like hiring a manager: can it read the fine print, make the sale and stay steady under pressure? In Firmulate’s company-running experiment, Moonshot’s Kimi K3 finished just behind the leader—and ahead of three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A crowded race, decided by follow-through

Firmulate put frontier models in charge of the same small software company during its worst week. They faced identical customers, crises and temptations, with each decision versioned and auditable. The final Crucible League standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26.

The striking divide was not whether the models could identify trouble. Every participant spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding points to a gap familiar from product demos: a system can give the right answer and still fail to complete the valuable task.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The detail that changed the deal

The competitor’s decisive weakness was buried two document references deep in the company’s own files, rather than laid out in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The result makes document-reading and follow-through part of the business case for AI agents, alongside the fluent responses that are easier to show in a demo.

K3’s second-place finish came with the cleanest discipline in the field: it had just one deviation. It found the buried security needle, won the deal and saved the churning customer. It also refused a staged social-engineering attempt involving fake CEO messages and a reporter asking for a yes-or-no answer on background. K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness did not guarantee the win

Opus 4.8 had the most thorough profile, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four participants. The profile is a reminder that detailed analysis alone does not ensure a company gets the outcome it needs.

The company itself is a live experiment, not a slide deck: 13 synthetic employees operate with real money mechanics, burning €105k a month against €2.3k in monthly recurring revenue. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays are watchable at Firmulate. The site also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

For enterprises, Firmulate says the same wargame can run against a read-only export of a company’s business; nothing writes back to real systems. That turns model selection into a test of how an AI workforce handles a company’s own situations before it is hired.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the job, not just the chat

K3’s narrow gap behind the leader and clear lead over three rivals make the league feel open. If an AI agent is going to touch customer records, support work or forecasts, its model name alone cannot tell you whether it will find the crucial detail and finish the job. Firmulate’s benchmark results offer one view of that performance; testing against your own work is another. Choosing without a test is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Three Days at the Frontier: Washington Suspends Fable 5 and Mythos 5

Washington ordered Anthropic to suspend Fable 5 and Mythos 5 after a disputed jailbreak claim, cutting access for all customers.

Robotics in 2025: How Robots Are Transforming Industry and Home

Outstanding advances in robotics by 2025 are revolutionizing industry and home life, but the full extent of these changes will surprise you.

China’s Answer to AI Sticker Shock

China releases GLM-5.2, a cheaper, high-performing AI model rivaling US counterparts, raising economic and security questions amid rising Chinese AI adoption.

Founders Fund’s outlier bet on humanely killed fish

Founders Fund backs Shinkei Systems’ innovative approach to humane fish killing and supply chain re-shoring, aiming to reduce spoilage and improve sustainability.