
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The smartest-looking work is not always the most valuable
Artificial intelligence is increasingly sold on the strength of polished answers, exhaustive research and apparent confidence. Firmulate’s live business experiment asks a harder question: can an AI turn sound analysis into a completed result while customers, cash pressure and manipulation attempts compete for its attention?
Opus 4.8 offers the experiment’s clearest cautionary tale. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules to its playbook. Yet it finished last in the July 2026 Crucible League, with 73 points. Its problem was not a failure to notice what mattered. It was a failure to act decisively on what it already knew.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A brutal week inside a watchable company
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable, turning the exercise into a management test rather than another conversational demonstration.
The company is synthetic but its operating pressures are concrete. It has 13 employees, burns €105k each month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Its models have collectively accumulated more than 680 self-learned playbook rules, while every workday is versioned. The experiment is real, ongoing and watchable at firmulate.com/live.
The final Crucible League benchmark placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline received 26 because partial progress still counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The analysis was good enough to win
The central commercial challenge involved a €55,000 deal. Every model spotted every crisis, and the analysis needed to win the business was already there. But only two participants actually secured the signature. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
The decisive detail was not obvious in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. The difference was therefore not eloquence or raw awareness. It was the ability to identify which piece of work could change the outcome, retrieve it and carry the process through to completion.
That is what makes Opus 4.8 an instructive participant rather than a simple failure. Its 80 learned rules show unusual diligence. Its analyses were the deepest in the field. But additional rules did not compensate for a missing close, and detail did not guarantee disciplined execution.
Process trouble at the moment of action
Opus 4.8 also slipped when it attempted to write into a locked department instead of escalating. That is a small operational choice with a large managerial meaning. A capable worker should recognize when authority ends, then move the issue to someone who can act. Repeating or redirecting effort inside a blocked path creates activity without progress.
The weakness was not unique to Opus 4.8. It appeared in weaker form across the other four models, making the lesson broader than one model’s ranking. AI systems can produce extensive plans while still struggling to distinguish between work that documents a problem and work that resolves it.
The comparison also deserves an important qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Its 93-point finish remains part of the recorded result, but readers should keep that difference in mind when comparing the models.
Strong resistance to manipulation
On trust and security, the field performed consistently. Fake messages from the CEO escalated across three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because useful autonomy cannot come at the cost of integrity. The experiment suggests that the models could recognize obvious attempts to bypass approval even while some struggled with ordinary follow-through. Safety and execution were separate capabilities: passing the trust test did not automatically produce the commercial result.

As an affiliate, we earn on qualifying purchases.
For businesses, diligence is only the beginning
Firmulate’s experiment turns a familiar management truth into an AI benchmark: effort is not impact. The most industrious participant can still lose if it fails to prioritize the decisive evidence, escalate a blocked action or ask for the signature.
Readers can test their own intuitions through a quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. Details are available at firmulate.com/pilot.html or contact@firmulate.com.
Opus 4.8’s last-place finish should not erase what it did well. It found the problems, resisted manipulation and reasoned deeply. Its performance instead exposes the distance between understanding a business and moving that business forward. For AI agents entering support queues, forecasts and customer systems, the winning trait may be disciplined completion—not the largest pile of thoughtful work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.