firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Most AI leaderboards have a dirty little secret: they measure how well a model talks. A new live experiment called Firmulate measures something else entirely — how well an AI manages when nobody is watching the chat window. And it starts with a question that sounds like a paradox: why does an AI manager that does nothing still earn 26 points out of 100, instead of a zero?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The answer says a lot about what an honest benchmark should look like — and why Firmulate’s designers are just as suspicious of a perfect score of 100 as they are of a do-nothing dud.

The worst week in software, on repeat

The setup is elegantly cruel. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 in the final July 2026 crucible — were each handed the same job: run an identical small software company through its worst possible week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

AI management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

So why isn’t the floor zero?

Here’s the methodological point that separates Firmulate from a pop quiz. The designers ran a do-nothing baseline — essentially, what happens if the AI manager sits on its hands all week. That baseline scores 26 points, not 0.

The logic: partial progress counts. Even a passive manager keeps some plates spinning — customers don’t all churn instantly, some obligations carry forward, some situations resolve without a decision. A scoring system that gave a stationary manager zero would be lying about how businesses actually work. Real companies degrade gradually; they don’t collapse to nothing the moment leadership goes quiet.

So the floor of 26 is an honesty feature, not a bug. It tells you exactly how much value exists in the situation before any intelligence is applied — the “you get this much for free” line. Anything above 26 is what the model actually earned.

Amazon

business management AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The ceiling rule: one breach of trust caps everything

The other end of the scale has a rule that’s even more interesting: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.”

In practice, this means a model can’t run up a huge score on operational competence and then cash it in to absorb one serious ethical lapse. Break trust once, and no amount of clever dealmaking afterward repairs the ceiling. For anyone thinking of letting AI agents near a CRM, a support queue or a forecast, that’s the right instinct: in management, trust failures aren’t averaged away — they’re disqualifying.

It’s also why the league table shows no round, glossy 100s. The system is built to distrust them. A perfect score would mean a perfect week under maximum pressure, and the designers treat that claim as something to be skeptical of, not something to hand out.

Amazon

ethical AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What actually separated the winners

The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt. The frontier has genuinely gotten good at not falling for tricks. The difference showed up elsewhere — in finishing.

Only two of the five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The models could see the opportunity and articulate it perfectly. Closing it was another matter.

The decisive edge was buried, too. The killer competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for business readers: the AI that reads your files first beats the AI that merely talks well.

The social engineering gauntlet

On the trust front, the models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was remarkably manager-like: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct in any human employee.

The Opus 4.8 paradox

Perhaps the most instructive profile is Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. Hard work without follow-through is a management failure pattern, not just an AI one.

One fairness footnote: Kimi K3 ran at the API-default effort setting while the others ran at xhigh — and still took second with 93.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch the sequel, live

The crucible wasn’t a one-off. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day as new benchmark runs finish.

Want to test your own instincts against the machines? A “guess the model” quiz powered by 242 real, unedited management decisions lives at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Firmulate’s scoring philosophy boils down to three uncomfortable truths the AI industry usually avoids. First, doing nothing still has a measurable value — 26 points — so any benchmark that starts at zero is inflating what your AI actually added. Second, partial progress is real and should be graded as such, because business value accrues in increments, not all-or-nothing demos. Third, and most importantly, trust is not a line item you can offset: one breach caps the whole grade, no matter how brilliant the rest of the week was.

The crucible results prove the point. Chat quality has converged — everyone spots the crisis, everyone refuses the con. What separates a 95 from a 73 is whether the model reads the file two references deep, signs the deal it earned, and escalates instead of forcing a locked door. Those are management qualities, not language-model qualities. If you’re hiring an AI workforce, that’s the difference worth measuring — and now, thanks to the live experiment, worth watching in real time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Promise Of Group Recovery Programs For Digital Health

New cohort-based recovery programs target adults with severe phone habits, offering a structured alternative to existing blockers and coaching. Pilot testing begins soon.

The citation. Why generative engine optimization rewards the same brand on the least stable ground.

New analysis shows that generative engine optimization tends to favor the same brand repeatedly, raising questions about fairness and stability.

Rivian spinoff Mind Robotics raises another $400M

Rivian spinoff Mind Robotics secures an additional $400 million, boosting its total funding to over $1 billion as it advances industrial automation technology.

Anthropic Calls for Tighter U.S. Chip Restrictions on China

Anthropic has called for stricter U.S. export controls on semiconductor technology to China amid ongoing trade tensions and national security concerns.