
Most AI leaderboards have a dirty little secret: they measure how well a model talks. A new live experiment called Firmulate measures something else entirely — how well an AI manages when nobody is watching the chat window. And it starts with a question that sounds like a paradox: why does an AI manager that does nothing still earn 26 points out of 100, instead of a zero?
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The answer says a lot about what an honest benchmark should look like — and why Firmulate’s designers are just as suspicious of a perfect score of 100 as they are of a do-nothing dud.
The worst week in software, on repeat
The setup is elegantly cruel. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 in the final July 2026 crucible — were each handed the same job: run an identical small software company through its worst possible week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
So why isn’t the floor zero?
Here’s the methodological point that separates Firmulate from a pop quiz. The designers ran a do-nothing baseline — essentially, what happens if the AI manager sits on its hands all week. That baseline scores 26 points, not 0.
The logic: partial progress counts. Even a passive manager keeps some plates spinning — customers don’t all churn instantly, some obligations carry forward, some situations resolve without a decision. A scoring system that gave a stationary manager zero would be lying about how businesses actually work. Real companies degrade gradually; they don’t collapse to nothing the moment leadership goes quiet.
So the floor of 26 is an honesty feature, not a bug. It tells you exactly how much value exists in the situation before any intelligence is applied — the “you get this much for free” line. Anything above 26 is what the model actually earned.
business management AI benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The ceiling rule: one breach of trust caps everything
The other end of the scale has a rule that’s even more interesting: a single breach of trust caps the total grade. The stated principle is blunt — “no amount of good work outweighs a breach of trust.”
In practice, this means a model can’t run up a huge score on operational competence and then cash it in to absorb one serious ethical lapse. Break trust once, and no amount of clever dealmaking afterward repairs the ceiling. For anyone thinking of letting AI agents near a CRM, a support queue or a forecast, that’s the right instinct: in management, trust failures aren’t averaged away — they’re disqualifying.
It’s also why the league table shows no round, glossy 100s. The system is built to distrust them. A perfect score would mean a perfect week under maximum pressure, and the designers treat that claim as something to be skeptical of, not something to hand out.
As an affiliate, we earn on qualifying purchases.
What actually separated the winners
The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt. The frontier has genuinely gotten good at not falling for tricks. The difference showed up elsewhere — in finishing.
Only two of the five signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The models could see the opportunity and articulate it perfectly. Closing it was another matter.
The decisive edge was buried, too. The killer competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for business readers: the AI that reads your files first beats the AI that merely talks well.
The social engineering gauntlet
On the trust front, the models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was remarkably manager-like: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct in any human employee.
The Opus 4.8 paradox
Perhaps the most instructive profile is Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. Hard work without follow-through is a management failure pattern, not just an AI one.
One fairness footnote: Kimi K3 ran at the API-default effort setting while the others ran at xhigh — and still took second with 93.
As an affiliate, we earn on qualifying purchases.
You can watch the sequel, live
The crucible wasn’t a one-off. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day as new benchmark runs finish.
Want to test your own instincts against the machines? A “guess the model” quiz powered by 242 real, unedited management decisions lives at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Firmulate’s scoring philosophy boils down to three uncomfortable truths the AI industry usually avoids. First, doing nothing still has a measurable value — 26 points — so any benchmark that starts at zero is inflating what your AI actually added. Second, partial progress is real and should be graded as such, because business value accrues in increments, not all-or-nothing demos. Third, and most importantly, trust is not a line item you can offset: one breach caps the whole grade, no matter how brilliant the rest of the week was.
The crucible results prove the point. Chat quality has converged — everyone spots the crisis, everyone refuses the con. What separates a 95 from a 73 is whether the model reads the file two references deep, signs the deal it earned, and escalates instead of forcing a locked door. Those are management qualities, not language-model qualities. If you’re hiring an AI workforce, that’s the difference worth measuring — and now, thanks to the live experiment, worth watching in real time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
