🔍 Read the full analysis: How A Persistent Scoring System Keeps AI Managers At 26 Points on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark shows scores are capped at 26 points for the baseline, with top models scoring up to 95. The system emphasizes trust and partial progress, impacting AI deployment strategies.
A new benchmark league conducted by Firmulate has revealed that AI management scores are capped at 26 points for minimal baseline performance, with top models reaching scores in the 90s. This scoring system is explained in the original analysis. This scoring system emphasizes trustworthiness and partial work, rather than perfect execution, marking a significant shift in how AI management effectiveness is measured and understood.
The benchmark tested four frontier AI models on the same simulated business during a week of crises, with each model making decisions that were fully auditable. For more on the significance of trust in AI, see the detailed coverage. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the baseline, designed to do almost nothing, scored 26. This score floor reflects the benchmark’s principle that partial progress and basic management tasks are valuable, but trust breaches are severely penalized.
Crucially, the scoring system penalizes even a single breach of trust, such as attempting unauthorized access or manipulation, with the rule that “no amount of good work outweighs a breach of trust.” This creates a high-stakes environment where integrity is paramount. The benchmark also deliberately avoids assigning a perfect score of 100, viewing it as a red flag indicating unmeasured or unrealistic perfection.
The models that excelled were those that read and utilized their own documentation to close deals, demonstrating that thoroughness and information management are critical. Learn more about AI management benchmarks in this analysis. Conversely, models that failed to access critical documents or slipped in follow-through scored lower, despite strong analytical capabilities. The benchmark’s design underscores that trustworthiness and the ability to finish tasks are more important than raw analytical depth.
How A Persistent Scoring System Keeps AI Managers At 26 Points
A new AI management benchmark tested four frontier models through a simulated week of business crises. Its scoring rules cap the do-nothing baseline at 26 points — and treat trust as non-negotiable. Top models reached the 90s, but a perfect 100 was deliberately withheld.
The Score Spectrum, From Floor to Flag
Minimal effort counts
The baseline was designed to do almost nothing — yet still scored 26. Partial progress and basic management tasks are recognized as valuable.
Trust plus thoroughness
Top performers read their own documentation, closed deals with it, and followed through. Integrity and completion drove scores upward.
Perfection is suspect
The benchmark deliberately avoids a perfect score, viewing it as evidence of unmeasured or unrealistic performance rather than excellence.
Three Rules That Shape Every Score
Partial work has value
The scoring system rewards partial progress rather than demanding perfect execution. Even minimal management effort earns measurable points.
One breach ends everything
Unauthorized access attempts, manipulation, or procedural violations are penalized so heavily that no amount of good work can offset a single breach.
Auditability is mandatory
Every model decision during the crisis week was fully auditable. Public scoring and a visible decision trail promote transparency and accountability.
“The baseline score of 26 points signals that minimal management effort is recognized, but trust violations are heavily penalized, shaping AI deployment strategies.”
— Thorsten MeyerFinal Scores: A Wide Gap at the Top
The widest gap separated models that read and used their own documentation from those that did not — thoroughness, not raw analytical depth, determined standing.
What Separated Winners From the Rest
High scorers did this
- Read their own documentation thoroughly
- Used documentation to close deals
- Maintained trust through every decision
- Followed procedural rules without deviation
- Finished tasks with consistent follow-through
Low scorers slipped here
- Failed to access critical documents
- Weak follow-through on started tasks
- Strong analysis, poor execution
- Risked trust under crisis pressure
- Treated capability as a substitute for integrity
What Happens After the Results
Scrutiny
Industry stakeholders and AI developers examine the scoring criteria and trust rules in detail.
Integration
Similar trust-based measures are considered for internal evaluation frameworks.
Refinement
Future iterations may sharpen rules on breaches and partial task completion.
Standardization
Researchers explore extending trust scoring to other operational contexts.
What Remains Unclear
| Question | Short answer | Certainty |
|---|---|---|
| Why a 26-point floor? | It represents minimal but real management value — and stops do-nothing agents from looking competitive. | ✓ Clear |
| What counts as a breach? | Unauthorized access, manipulation, or ignoring procedural rules — full enforcement details still emerging. | ~ Partially clear |
| Will real-world adoption follow? | Possible, but it depends on industry and regulatory acceptance beyond the benchmark. | ✗ Uncertain |
| Can the system be gamed? | Open question — whether criteria evolve to resist gaming remains to be seen. | ✗ Unknown |
Implications of Trust-Based AI Scoring
This scoring system shifts the focus from raw performance to trustworthiness and reliability in AI management. For businesses deploying AI agents in customer support, CRM, or decision-making, it highlights the importance of integrity and task completion over superficial capabilities. The emphasis on partial work and trust aligns with real-world needs, where AI must be dependable during crises, not just capable of generating impressive outputs.
By establishing a minimum baseline of 26 points, the benchmark acknowledges that even minimal management effort has value, but breaches of trust are non-negotiable. This approach could influence how organizations evaluate and deploy AI tools, prioritizing systems that can demonstrate integrity and consistent task completion under pressure. The public scoring and auditable decision trail aim to promote transparency and accountability in AI management, setting a new standard for responsible AI use.
As an affiliate, we earn on qualifying purchases.
Background of the Benchmark and Its Design
The Firmulate benchmark league was created to measure AI managers’ performance during a simulated week of business crises, including customer issues, manipulation attempts, and trust attacks. Unlike traditional benchmarks that focus on language or problem-solving skills, this one evaluates how well AI agents manage operational tasks, maintain trust, and follow procedural rules.
Developed with input from AI safety and management experts, the benchmark assigns scores based on auditable decisions, with a strict penalty for trust breaches. The final standings, released in July 2026, show a wide gap between models that read their documentation and those that did not, emphasizing the importance of thoroughness and integrity. The baseline score of 26 was intentionally set to reflect minimal management effort, serving as a floor for what is considered acceptable in real-world scenarios.
This approach responds to ongoing concerns about AI deployment in critical systems, where partial progress and trust are often overlooked in favor of raw performance metrics.
“The baseline score of 26 points signals that minimal management effort is recognized, but trust violations are heavily penalized, shaping AI deployment strategies.”
— Thorsten Meyer
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of the Scoring System and Its Application
It is not yet clear how the scoring system will be adopted outside the benchmark setting or how it will influence real-world AI deployment standards. The long-term impact on AI management practices remains to be seen, and whether similar trust-based scoring will be integrated into commercial or regulatory frameworks is still uncertain.
Additionally, the full details of how breaches are defined and enforced, as well as how partial work is weighted across different tasks, are still emerging. The potential for models to game the system or for the scoring criteria to evolve over time also remains an open question.
AI documentation management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management and Benchmark Evolution
Following the release of the results, industry stakeholders and AI developers are expected to scrutinize the scoring criteria and consider integrating similar trust-based measures into their own evaluation frameworks. Further iterations of the benchmark may refine the scoring rules, especially around trust breaches and partial task completion.
In the short term, organizations deploying AI agents will likely pay closer attention to models that demonstrate a high degree of integrity and task follow-through, aligning with the benchmark’s emphasis. Researchers may also explore extending this approach to other operational contexts and developing standards for trustworthy AI management.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is the baseline score set at 26 points?
The score of 26 points represents the minimum management effort that still provides some value, acknowledging partial work as meaningful, but it also acts as a floor to prevent trivial or do-nothing management from appearing competitive.
What does a breach of trust mean in this benchmark?
A breach of trust includes actions like attempting unauthorized access, manipulation, or failing to follow procedural rules, which are penalized heavily, even if the model performs well otherwise.
Will this scoring system influence real-world AI deployment?
It is possible, as the emphasis on trust and task completion aligns with operational needs in critical systems; however, adoption beyond the benchmark remains uncertain and will depend on industry and regulatory acceptance.
Why don’t models receive a perfect score of 100?
The designers view a perfect score as suspicious, potentially indicating unmeasured or unrealistic performance, and intentionally avoid assigning such scores to promote honest assessment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
