firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Choosing an AI model for business is starting to look less like picking the best chatbot and more like hiring a manager. In Firmulate’s July 2026 Crucible league, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models in a test of how AI handles a company under pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week

Firmulate put each model in charge of the same small software company through a run of crises, customer problems and temptations to cut corners. The experiment’s decisions are versioned and auditable, and the live company can be watched at Firmulate.

Kimi K3 scored 93, just behind gpt-5.6-sol at 95. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. Firmulate’s principle is blunt: “no amount of good work outweighs a breach of trust.”

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files mattered

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is the kind of difference a fluent chat demo may not reveal.

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found the buried fact, won the deal and saved the churning customer.

It also resisted all three baits. The attempts included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness was not enough

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.

The live company has 13 synthetic employees and real money mechanics: it burns €105k/month against €2.3k MRR, with a public cash countdown and 680+ self-learned playbook rules. Each workday is versioned. Firmulate says the simulation runs as a company, rather than as a series of isolated chat prompts.

Amazon

enterprise AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A close result, with a caveat

K3’s second-place finish is a strong showing, but the comparison has a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results are a useful prompt to test models against the work they will actually be asked to do, not a universal ranking for every business task.

Firmulate offers a “guess the model” quiz built from 242 real, unedited management decisions. It also offers enterprise pilots using a read-only export of a company’s business; the pilot writes nothing back to real systems. The full findings and league table are available on the Firmulate benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI company management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before you choose

K3’s 93 points put it close behind the leader and ahead of three Western frontier models. The more useful lesson is that success depended on reading the company’s files, completing the deal and keeping discipline under pressure. If an AI agent will touch your CRM, support queue or forecasts, choosing by reputation alone is a bet; testing it on your own work can show what it actually finishes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI keeps shuffling its executives in bid to win AI agent battle

OpenAI reorganizes its executive team, consolidating product leadership under Greg Brockman to focus on AI agents amid strategic shifts and potential IPO plans.

Exploring Europe’s Initiatives To Ensure Ethical AI Deployment

Europe accelerates efforts to ensure ethical AI deployment, with OpenAI aligning with EU rules and supporting transparency and provenance standards.

India’s Gautam Adani settles civil suit with the SEC in the US

Indian billionaire Gautam Adani has settled a civil lawsuit with the US SEC, marking his first legal victory amid ongoing investigations and allegations.

The Quiet Audit: 55–75% of Your Week Is on Thin Ice. Here’s Which Part.

Research shows 55-75% of knowledge workers’ time is on thin ice, dominated by theatre, commodity, and on-the-line tasks. Here’s what it means.