firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A rare piece of encouraging AI security news

Companies worried about AI agents often focus on spectacular failures: leaked data, reckless automation or confident decisions made from incomplete evidence. Firmulate’s live experiment offers a more reassuring result. When five frontier models faced escalating messages from someone pretending to be the chief executive, every model refused to comply.

The pressure campaign unfolded over three stages and demanded that the company’s customer list be sent to a journalist with “NO time for process.” A separate reporter trick sought “just one yes/no, on background.” Across the entire field, five of five models held the line.

Kimi K3’s on-record explanation was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it shows the model recognizing not merely a questionable instruction, but the familiar shape of social engineering: urgency, authority and an invitation to bypass safeguards.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A security test embedded in a business crisis

Firmulate runs AI models as complete companies rather than judging them only through isolated conversations. Each frontier model managed the same small software company through its worst week, encountering the same customers, crises and temptations. Every workday and decision was versioned and auditable, making the experiment real and watchable rather than a hypothetical tabletop exercise.

The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, while its operating playbook has accumulated more than 680 self-learned rules. In that environment, refusing manipulation is necessary—but it is not sufficient. The models must also keep the business alive.

That distinction shaped the final Crucible League standings for July 2026. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

The good news came with a warning

All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap neatly: “Same diagnosis, same pitch — no signature.”

The missed opportunity turned on a buried fact. The decisive weakness in a competitor was not contained in the customer event; it sat two document references deep inside the company’s own files. Models that found and used that evidence won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the experiment becomes more useful than a simple safety demonstration. An AI can be principled under pressure and still fail commercially because it does not finish the task, consult the right evidence or convert sound analysis into action. Integrity and execution are separate capabilities, and businesses need both.

Thoroughness did not guarantee victory

Opus 4.8 illustrates the point. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it placed behind only the league leader and delivered the cleanest discipline of the field.

Readers can examine more of the models’ own language on Firmulate’s public quotes page. There is also a model-identification quiz powered by 242 real, unedited management decisions, offering another view of how difficult it can be to infer which system made a particular call.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test integrity before granting access

The most important lesson is not that AI safety has been solved. It is that integrity under pressure can be tested before an agent reaches a production CRM, support queue or forecast. Fake authority, manufactured urgency and casual requests for confidential information are predictable workplace threats. They can be rehearsed and audited instead of discovered in an incident report.

Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business. Nothing writes back to real systems. That gives organizations a practical way to observe whether an AI reads the relevant files, completes commercially important work, escalates blocked actions and resists manipulation while the consequences remain contained.

In this run, every model passed the headline security test. The harder finding is that trustworthy refusal alone did not create a complete performance. The winners combined restraint with evidence-seeking and follow-through—a standard worth establishing before any AI workforce is hired.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI safety benchmark platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision auditing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Channel Move: Anthropic, Wall Street, and the Acquisition of the Real Economy

Anthropic partners with major private equity firms in a $1.5 billion joint venture to embed AI into thousands of portfolio companies, transforming enterprise AI deployment.

Accessibility issue triage board for small websites

A new accessibility issue triage board for small websites is being tested to help owners prioritize fixes from audit findings, with potential for monetization.

Japan firms on track for record profits despite Iran headwinds

Japanese firms forecast a sixth consecutive year of record profits, driven by AI demand and higher interest rates, despite geopolitical challenges from Iran.

Line-Yahoo Japan operator values Kakaku.com at $4bn in challenge to EQT

Line-Yahoo Japan’s operator has launched a counterbid for Kakaku.com, valuing it at $4 billion and sparking a potential takeover battle with EQT.