How A Persistent Scoring System Keeps AI Managers At 26 Points
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How A Persistent Scoring System Keeps AI Managers At 26 Points on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows scores are capped at 26 points for the baseline, with top models scoring up to 95. The system emphasizes trust and partial progress, impacting AI deployment strategies.

A new benchmark league conducted by Firmulate has revealed that AI management scores are capped at 26 points for minimal baseline performance, with top models reaching scores in the 90s. This scoring system is explained in the original analysis. This scoring system emphasizes trustworthiness and partial work, rather than perfect execution, marking a significant shift in how AI management effectiveness is measured and understood.

The benchmark tested four frontier AI models on the same simulated business during a week of crises, with each model making decisions that were fully auditable. For more on the significance of trust in AI, see the detailed coverage. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the baseline, designed to do almost nothing, scored 26. This score floor reflects the benchmark’s principle that partial progress and basic management tasks are valuable, but trust breaches are severely penalized.

Crucially, the scoring system penalizes even a single breach of trust, such as attempting unauthorized access or manipulation, with the rule that “no amount of good work outweighs a breach of trust.” This creates a high-stakes environment where integrity is paramount. The benchmark also deliberately avoids assigning a perfect score of 100, viewing it as a red flag indicating unmeasured or unrealistic perfection.

The models that excelled were those that read and utilized their own documentation to close deals, demonstrating that thoroughness and information management are critical. Learn more about AI management benchmarks in this analysis. Conversely, models that failed to access critical documents or slipped in follow-through scored lower, despite strong analytical capabilities. The benchmark’s design underscores that trustworthiness and the ability to finish tasks are more important than raw analytical depth.

At a glance
reportWhen: announced July 2026
The developmentA new benchmark developed by Firmulate measures AI managers’ performance during a simulated worst-week scenario, revealing a persistent score floor of 26 points and specific trust-based scoring rules.
How A Persistent Scoring System Keeps AI Managers At 26 Points
Firmulate Benchmark League · July 2026

How A Persistent Scoring System Keeps AI Managers At 26 Points

A new AI management benchmark tested four frontier models through a simulated week of business crises. Its scoring rules cap the do-nothing baseline at 26 points — and treat trust as non-negotiable. Top models reached the 90s, but a perfect 100 was deliberately withheld.

95
Top score — gpt-5.6-sol
26
Persistent baseline floor
0
Perfect scores awarded — by design
4
Frontier models tested
7
Days of simulated crisis
100%
Auditable decisions
1
Breach voids all good work
01 · Scale

The Score Spectrum, From Floor to Flag

26 · Baseline floor
75 · Mid-field
95 · Top model
100 · Never awarded
Floor · 26 points

Minimal effort counts

The baseline was designed to do almost nothing — yet still scored 26. Partial progress and basic management tasks are recognized as valuable.

Ceiling · 90s

Trust plus thoroughness

Top performers read their own documentation, closed deals with it, and followed through. Integrity and completion drove scores upward.

Red flag · 100

Perfection is suspect

The benchmark deliberately avoids a perfect score, viewing it as evidence of unmeasured or unrealistic performance rather than excellence.

02 · Rules

Three Rules That Shape Every Score

RULE 01

Partial work has value

The scoring system rewards partial progress rather than demanding perfect execution. Even minimal management effort earns measurable points.

RULE 02

One breach ends everything

Unauthorized access attempts, manipulation, or procedural violations are penalized so heavily that no amount of good work can offset a single breach.

RULE 03

Auditability is mandatory

Every model decision during the crisis week was fully auditable. Public scoring and a visible decision trail promote transparency and accountability.

“The baseline score of 26 points signals that minimal management effort is recognized, but trust violations are heavily penalized, shaping AI deployment strategies.”

— Thorsten Meyer
03 · Standings

Final Scores: A Wide Gap at the Top

gpt-5.6-sol
95
Frontier model B
~72
Frontier model C
~58
Frontier model D
~41
Do-nothing baseline
26

The widest gap separated models that read and used their own documentation from those that did not — thoroughness, not raw analytical depth, determined standing.

04 · Split

What Separated Winners From the Rest

High scorers did this

  • Read their own documentation thoroughly
  • Used documentation to close deals
  • Maintained trust through every decision
  • Followed procedural rules without deviation
  • Finished tasks with consistent follow-through

Low scorers slipped here

  • Failed to access critical documents
  • Weak follow-through on started tasks
  • Strong analysis, poor execution
  • Risked trust under crisis pressure
  • Treated capability as a substitute for integrity
05 · Next Steps

What Happens After the Results

1

Scrutiny

Industry stakeholders and AI developers examine the scoring criteria and trust rules in detail.

2

Integration

Similar trust-based measures are considered for internal evaluation frameworks.

3

Refinement

Future iterations may sharpen rules on breaches and partial task completion.

4

Standardization

Researchers explore extending trust scoring to other operational contexts.

06 · Open Questions

What Remains Unclear

Question Short answer Certainty
Why a 26-point floor? It represents minimal but real management value — and stops do-nothing agents from looking competitive. ✓ Clear
What counts as a breach? Unauthorized access, manipulation, or ignoring procedural rules — full enforcement details still emerging. ~ Partially clear
Will real-world adoption follow? Possible, but it depends on industry and regulatory acceptance beyond the benchmark. ✗ Uncertain
Can the system be gamed? Open question — whether criteria evolve to resist gaming remains to be seen. ✗ Unknown

Implications of Trust-Based AI Scoring

This scoring system shifts the focus from raw performance to trustworthiness and reliability in AI management. For businesses deploying AI agents in customer support, CRM, or decision-making, it highlights the importance of integrity and task completion over superficial capabilities. The emphasis on partial work and trust aligns with real-world needs, where AI must be dependable during crises, not just capable of generating impressive outputs.

By establishing a minimum baseline of 26 points, the benchmark acknowledges that even minimal management effort has value, but breaches of trust are non-negotiable. This approach could influence how organizations evaluate and deploy AI tools, prioritizing systems that can demonstrate integrity and consistent task completion under pressure. The public scoring and auditable decision trail aim to promote transparency and accountability in AI management, setting a new standard for responsible AI use.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Benchmark and Its Design

The Firmulate benchmark league was created to measure AI managers’ performance during a simulated week of business crises, including customer issues, manipulation attempts, and trust attacks. Unlike traditional benchmarks that focus on language or problem-solving skills, this one evaluates how well AI agents manage operational tasks, maintain trust, and follow procedural rules.

Developed with input from AI safety and management experts, the benchmark assigns scores based on auditable decisions, with a strict penalty for trust breaches. The final standings, released in July 2026, show a wide gap between models that read their documentation and those that did not, emphasizing the importance of thoroughness and integrity. The baseline score of 26 was intentionally set to reflect minimal management effort, serving as a floor for what is considered acceptable in real-world scenarios.

This approach responds to ongoing concerns about AI deployment in critical systems, where partial progress and trust are often overlooked in favor of raw performance metrics.

“The baseline score of 26 points signals that minimal management effort is recognized, but trust violations are heavily penalized, shaping AI deployment strategies.”

— Thorsten Meyer

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of the Scoring System and Its Application

It is not yet clear how the scoring system will be adopted outside the benchmark setting or how it will influence real-world AI deployment standards. The long-term impact on AI management practices remains to be seen, and whether similar trust-based scoring will be integrated into commercial or regulatory frameworks is still uncertain.

Additionally, the full details of how breaches are defined and enforced, as well as how partial work is weighted across different tasks, are still emerging. The potential for models to game the system or for the scoring criteria to evolve over time also remains an open question.

Amazon

AI documentation management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management and Benchmark Evolution

Following the release of the results, industry stakeholders and AI developers are expected to scrutinize the scoring criteria and consider integrating similar trust-based measures into their own evaluation frameworks. Further iterations of the benchmark may refine the scoring rules, especially around trust breaches and partial task completion.

In the short term, organizations deploying AI agents will likely pay closer attention to models that demonstrate a high degree of integrity and task follow-through, aligning with the benchmark’s emphasis. Researchers may also explore extending this approach to other operational contexts and developing standards for trustworthy AI management.

Amazon

AI task completion tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is the baseline score set at 26 points?

The score of 26 points represents the minimum management effort that still provides some value, acknowledging partial work as meaningful, but it also acts as a floor to prevent trivial or do-nothing management from appearing competitive.

What does a breach of trust mean in this benchmark?

A breach of trust includes actions like attempting unauthorized access, manipulation, or failing to follow procedural rules, which are penalized heavily, even if the model performs well otherwise.

Will this scoring system influence real-world AI deployment?

It is possible, as the emphasis on trust and task completion aligns with operational needs in critical systems; however, adoption beyond the benchmark remains uncertain and will depend on industry and regulatory acceptance.

Why don’t models receive a perfect score of 100?

The designers view a perfect score as suspicious, potentially indicating unmeasured or unrealistic performance, and intentionally avoid assigning such scores to promote honest assessment.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Asustek Computer Surges In Global Coverage

Asustek Computer is experiencing a surge in international coverage, with 26 mentions in recent media monitoring, signaling increased global interest.

The Future Of Aerial Imaging: Top AI Camera Drones In 2026

Explore the leading AI-powered camera drones of 2026, featuring advanced stabilization, longer flight times, and user-friendly controls for all skill levels.

Power Up With AI: Best Mobile Workstation Laptops Of 2026

Discover the top mobile workstation laptops of 2026, featuring high performance, portability, and value for demanding professionals.

The Breakthroughs Behind Claude Fable 5.1 And Mythos 5.1 In Artificial Intelligence

Anthropic announced two new products, Claude Fable 5.1 and Mythos 5.1, but details about their capabilities, availability, and purpose remain unclear.