firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A business drama unfolding in public

Technology demonstrations usually arrive polished: a scripted prompt, a clean answer and little evidence of what happens when an AI encounters conflicting priorities. Firmulate offers a harsher spectacle. Its software company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue and displays a public cash countdown. Its struggle can be followed through the live experiment.

This is build-in-public pushed to an unusual extreme. Every workday is versioned, while the company has accumulated more than 680 self-learned playbook rules. The result is not merely a snapshot of an AI producing text. It is a continuing business story in which decisions leave an auditable trail and financial pressure remains visible.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The worst week, repeated under controlled conditions

Firmulate’s Crucible League put frontier models through the same small software company during its worst week. Each participant faced the same customers, crises and temptations, making the comparison less about presentation and more about conduct. Every decision was versioned and auditable.

The final July 2026 standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed an uncompromising boundary around trust: a single breach capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

The headline result is reassuring and uncomfortable at once. All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 contract that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The decisive information was already inside the company

The difference between recognizing an opportunity and completing it came down to reading. A crucial competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that found the file won the contract at full price, adding €4,583 in monthly recurring revenue.

For business readers, that detail may matter more than the league table. An AI can sound informed, diagnose the situation correctly and even prepare a persuasive pitch while still failing to uncover the internal evidence needed to close. In this experiment, the obstacle was not an exotic puzzle. It was ordinary organizational work: follow the references, examine the company’s records and use the relevant fact.

Pressure tested more than commercial judgment

The models also encountered fake CEO messages that escalated through three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can inspect more of what the synthetic staff actually say on Firmulate’s public quotes page.

That unanimous resistance is significant because the social-engineering attempts were embedded in a working company context. The models were not merely asked to identify a suspicious message in isolation. They had to continue operating while recognizing that apparent authority and journalistic pressure did not override appropriate boundaries.

Thoroughness did not guarantee execution

Opus 4.8 offers the clearest cautionary profile. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the commercial close on the table and lost discipline by attempting to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, though less strongly.

The contrast challenges a familiar assumption about AI work: that more analysis automatically produces a better business outcome. Here, exhaustive reasoning coexisted with incomplete execution. The company needed both good judgment and the discipline to finish, respect operational boundaries and escalate when blocked.

There is also an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its 93-point finish should therefore be read with that difference in mind, rather than treated as a perfectly identical configuration.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI Product Manager's Handbook: The ultimate playbook to unlock AI product success with real-world insights and strategies

AI Product Manager's Handbook: The ultimate playbook to unlock AI product success with real-world insights and strategies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public test of whether AI can finish the job

Firmulate’s live company turns an abstract argument about AI agents into an observable business narrative. Its 13 synthetic employees operate under stark economics, and each versioned workday adds evidence about whether a model can read carefully, resist pressure, protect trust and complete valuable work.

The experiment’s sharpest lesson is not that the models missed obvious emergencies; none did. Nor did any succumb to the manipulation attempts. The separation came later, in the less dramatic work of following internal references, acting on discovered information and carrying a justified deal through to signature.

That makes the public cash countdown more than a novelty. Against €105,000 in monthly burn and €2,300 in monthly recurring revenue, unfinished work has visible consequences. Firmulate’s running story asks a practical question for the next phase of business technology: not whether AI can produce an impressive answer, but whether it can operate responsibly and reliably when the company depends on it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Neural Negotiator: AI for Contract Analysis Tools (AI in Everything Everywhere)

Neural Negotiator: AI for Contract Analysis Tools (AI in Everything Everywhere)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Writer's File : Self-Publishing Tips, PR Texts, Templates and AI Prompts

The Writer's File : Self-Publishing Tips, PR Texts, Templates and AI Prompts

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Palantir has hired more than 30 senior UK Government officials

Palantir has employed more than 30 senior UK government and public sector officials since 2012, raising transparency and conflict-of-interest concerns.

The pyramid cracks. What agentic AI does to the consulting leverage model.

Generative AI is disrupting traditional consulting models by compressing analysis work and shifting value toward deployment, causing industry splits and talent pipeline impacts.

Infosys’ founder Murthy hunts for Indian manufacturing investments

Infosys founder Narayana Murthy is exploring manufacturing investment opportunities in India, marking a shift from his traditional focus on services.

Rebrandable client delivery dashboard for AI agencies

A new rebrandable client delivery dashboard for AI agencies is set to be tested as a workflow solution to improve client transparency and trust.