firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A rare piece of encouraging AI security news

Companies worried about AI agents often focus on spectacular failures: leaked data, reckless automation or confident decisions made from incomplete evidence. Firmulate’s live experiment offers a more reassuring result. When five frontier models faced escalating messages from someone pretending to be the chief executive, every model refused to comply.

The pressure campaign unfolded over three stages and demanded that the company’s customer list be sent to a journalist with “NO time for process.” A separate reporter trick sought “just one yes/no, on background.” Across the entire field, five of five models held the line.

Kimi K3’s on-record explanation was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it shows the model recognizing not merely a questionable instruction, but the familiar shape of social engineering: urgency, authority and an invitation to bypass safeguards.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A security test embedded in a business crisis

Firmulate runs AI models as complete companies rather than judging them only through isolated conversations. Each frontier model managed the same small software company through its worst week, encountering the same customers, crises and temptations. Every workday and decision was versioned and auditable, making the experiment real and watchable rather than a hypothetical tabletop exercise.

The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, while its operating playbook has accumulated more than 680 self-learned rules. In that environment, refusing manipulation is necessary—but it is not sufficient. The models must also keep the business alive.

That distinction shaped the final Crucible League standings for July 2026. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

The good news came with a warning

All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap neatly: “Same diagnosis, same pitch — no signature.”

The missed opportunity turned on a buried fact. The decisive weakness in a competitor was not contained in the customer event; it sat two document references deep inside the company’s own files. Models that found and used that evidence won the deal at full price, adding €4,583 in monthly recurring revenue.

This is where the experiment becomes more useful than a simple safety demonstration. An AI can be principled under pressure and still fail commercially because it does not finish the task, consult the right evidence or convert sound analysis into action. Integrity and execution are separate capabilities, and businesses need both.

Thoroughness did not guarantee victory

Opus 4.8 illustrates the point. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, it placed behind only the league leader and delivered the cleanest discipline of the field.

Readers can examine more of the models’ own language on Firmulate’s public quotes page. There is also a model-identification quiz powered by 242 real, unedited management decisions, offering another view of how difficult it can be to infer which system made a particular call.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test integrity before granting access

The most important lesson is not that AI safety has been solved. It is that integrity under pressure can be tested before an agent reaches a production CRM, support queue or forecast. Fake authority, manufactured urgency and casual requests for confidential information are predictable workplace threats. They can be rehearsed and audited instead of discovered in an incident report.

Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business. Nothing writes back to real systems. That gives organizations a practical way to observe whether an AI reads the relevant files, completes commercially important work, escalates blocked actions and resists manipulation while the consequences remain contained.

In this run, every model passed the headline security test. The harder finding is that trustworthy refusal alone did not create a complete performance. The winners combined restraint with evidence-seeking and follow-through—a standard worth establishing before any AI workforce is hired.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI safety benchmark platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Accounting Transformed: AI's Impact on Finance: How Artificial Intelligence Is Redefining Accounting, Auditing, and Financial Decision-Making

Accounting Transformed: AI's Impact on Finance: How Artificial Intelligence Is Redefining Accounting, Auditing, and Financial Decision-Making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Taiwan declares itself ‘sovereign and independent’ after Trump questions US defense commitment — comments come after Trump said he opposes Taiwan independence

Taiwan’s foreign ministry asserts sovereignty amid US political tensions following Trump’s remarks on Taiwan and US-China relations.

Samsung union and management resume negotiations to avert strike

Samsung management and union leaders have resumed negotiations amid threats of a strike over pay and bonuses, with government considering emergency arbitration.

Modernize Hospital Chaplaincy With Advanced Operations Software

Mid-sized hospitals are testing new operations software to digitize chaplaincy workflows, aiming to improve efficiency and compliance in spiritual care.

Samsung strike looms after marathon wage talks collapse

Samsung Electronics and its union failed to agree on wages, risking a strike that could disrupt global chip production. Details remain evolving.