firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A benchmark win is not a management credential

Technology buyers have learned to compare AI models through coding leaderboards and chat arenas. Those tests can reveal whether a model produces a strong answer, but they say little about what happens when an agent must triage competing crises, operate under capacity pressure and live with the consequences of yesterday’s decisions.

That gap becomes important when AI moves beyond the chat window. An agent touching a customer queue, revenue forecast or sales process is no longer merely answering questions. It is managing incomplete information, organizational boundaries and temptations to take shortcuts. The relevant category is management quality, not chat quality.

Firmulate, an AI company emulator, makes that distinction visible by giving frontier models responsibility for the same small software company during its worst week. Each faced the same customers, crises and temptations, with every decision versioned and auditable. The experiment is real, public and watchable.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between seeing and finishing

The final July 2026 Crucible League table put gpt-5.6-sol in front with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Yet a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Those results matter less as a horse race than as evidence of a management gap. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its most unsettling: “Same diagnosis, same pitch — no signature.” Recognition was universal; completion was not.

The decisive information was not sitting conveniently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. This is the sort of behavior conventional demonstrations can obscure. A polished response may look intelligent while leaving the crucial document unread and the commercial result unrealized.

Pressure tests character as well as competence

Firmulate also subjected the models to fake CEO messages escalating over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest security posture: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result deserves attention because enterprise agents will encounter requests that sound urgent, authoritative and plausible. The consequential question is not merely whether a model recognizes manipulation in isolation. It is whether the model preserves trust while customers, cash and internal demands are all competing for attention.

The company makes those trade-offs unusually tangible. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its cash countdown is public, its workdays are versioned and its staff have accumulated 680+ self-learned playbook rules. This creates consequences across days rather than treating each prompt as an independent performance.

Thoroughness can still lose

Opus 4.8 illustrates why answer quality and management quality cannot be treated as synonyms. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared, less strongly, in all four other participants.

This is not an argument against analysis. It is an argument that analysis is only one part of responsible execution. A manager must discover the relevant fact, choose a course, respect organizational controls and carry the work through to a legitimate conclusion. Detailed reasoning that never becomes an authorized action can still be commercially useless.

There is also an important fairness note: K3 ran at its API default without an effort parameter, while the others ran at xhigh. Readers comparing the full benchmark results should keep that difference in view rather than treating the standings as context-free measurements.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scenario names may be the next curriculum

Churn wave, price increase, downround and PR crisis are more revealing labels for agent evaluation than another abstract reasoning category. They describe situations in which priorities collide and an apparently minor omission can affect customers, revenue or board trust.

Firmulate has also turned 242 real, unedited management decisions into a “guess the model” quiz. The point is not simply to identify a writing style. It is to confront how difficult it can be to distinguish confident language from dependable judgment.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That offers a practical bridge between public benchmarks and deployment: test an AI workforce against recognizable files, pressures and consequences before granting it operational authority.

Coding scores and chat preferences will remain useful. But once agents are expected to manage work, the decisive test changes. Can they read deeply, resist manipulation, respect boundaries, stay honest toward the board and finish what they start? Firmulate’s league suggests that the distance between a good answer and a good manager is still large enough to cost the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will Vietnam’s Asia pivot pay off?

Vietnam’s recent diplomatic outreach signals a strategic pivot towards Asia. Experts debate whether this move will succeed amid US-China tensions.

The Evolution Of Fintech Through Artificial Intelligence

Fintech sector collapses in 2022-24, then pivots to AI-driven infrastructure, with funding shifting to agentic payments and machine-initiated commerce.

Michael O’Leary in line for €150mn payout in latest Ryanair contract

Ryanair’s CEO Michael O’Leary is reportedly in line for a €150 million payout under his latest contract, sparking discussions on executive compensation.

Data Centers Surges In Global Coverage

Data center mentions increase significantly worldwide, with GDELT recording 16 times the baseline in recent monitoring, highlighting rapid expansion.