Can Human Review Keep Pace With AI Production?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can Human Review Keep Pace With AI Production? on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI published 722 mathematical manuscripts produced through a programme that explored about 4,000 problems, while the source report describes expert review of an earlier result by five leading mathematicians. Software and professional-workflow figures point to a similar challenge: AI can increase output faster than people can check it. The scale and consequences vary by field, and some supporting data comes from companies that sell review tools.

OpenAI published 722 mathematical manuscripts this week from a programme that explored about 4,000 problems, bringing renewed attention to a growing mismatch: AI systems can produce work quickly, but checking whether it is correct and useful still relies heavily on people. The development matters beyond mathematics, as software and professional-workflow data cited in the source also point to review becoming a constraint on using AI-generated output.

The manuscripts were grouped into 372 families of results, and the average result took about three hours of compute to produce, according to the source report. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that unformalized results “could have issues.” The report says an earlier result from the same programme, a counterexample to an old Erdős conjecture, received careful verification from five leading mathematicians. That example illustrates the human effort involved; it does not establish how every manuscript was reviewed.

Software data in the report suggests the same tension, though the figures come from different studies and should not be treated as directly comparable. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes.

A peer-reviewed 2026 study cited by the source found 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with zero review during high-adoption periods. Several cited software-data providers sell code-review tools, a commercial interest that warrants care when interpreting their results. The source also describes an Ironclad partnership involving OpenAI’s GPT-6 Astra: across 11 contracting tasks, the model met 55% of evaluation criteria on average. That is a reported evaluation result, not evidence that the remaining criteria were always missed in live contracts.

At a glance
reportWhen: Reported this week; software and workfl…
The developmentOpenAI’s publication of 722 AI-produced mathematics manuscripts has renewed attention on whether human reviewers can keep up with rapidly growing volumes of AI-generated work.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Is Becoming a Constraint

The central consequence is operational: organisations may be able to generate more drafts, proofs, code changes and contract work than their experts can responsibly approve. If review capacity does not expand, teams face trade-offs among delays, shallow checks and unreviewed output. The cited studies show these patterns in software, but they do not establish a single economy-wide rate or prove that every organisation is experiencing the same effects.

Reviewers also carry responsibility that an AI system does not. A person signs a contract, approves an engineering design or answers for a published result. That makes expert judgment a potential bottleneck even where automated checks can catch some errors. The report’s interpretation is that demand may rise for people capable of validating AI-assisted work; whether that produces a broad wage or employment premium is not demonstrated by the figures provided.

There is also a training concern. Experienced reviewers typically build judgment by doing the underlying work: writing code, drafting contracts or developing mathematical arguments. If AI takes over too much junior work, future professionals may have fewer chances to acquire the skills needed to review it. The scale of that risk is uncertain, but it raises a practical question for employers: how to gain productivity while preserving routes for workers to develop expertise.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workplaces

The source report draws examples from mathematics, software and contract work, but they measure different things. In mathematics, formal verification can establish that a proof follows from its stated assumptions and conclusion. It does not by itself establish that the theorem addresses the right question or that the result is significant. Human specialists still make those judgments.

In software, pull-request data tracks review waits, acceptance and whether a human looked at a change. Those indicators provide a view of workflow, but they do not alone show the quality or downstream impact of every change. The figures cited come from separate analyses with different methods, populations and periods. The report notes that some sources sell code-review products, so their findings should be read with that commercial context in mind.

The contracting example offers a separate measure: performance against criteria on 11 tasks. It points to a gap between model capability and the evaluation standard used, not a general measure of legal accuracy. Across all three cases, the useful distinction is between producing an answer and validating its fitness for use.

“Five of the world’s leading mathematicians carefully verified an earlier counterexample to an old Erdős conjecture.”

— Source report’s description of the mathematics review

Amazon

proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The available figures do not provide a single, consistent measure of review quality across fields. It is unclear how many of the 722 mathematical manuscripts were independently checked, how review was allocated among them, or how many results later required correction. OpenAI’s warning concerns unformalized work, but the source does not specify the share of manuscripts in that category.

The software statistics also leave important questions open: what counts as a review, how studies account for differences in task difficulty, and whether lower acceptance rates reflect errors, stricter scrutiny or other factors. Some sources have commercial interests in review tools, while the cited peer-reviewed study is a separate source. The Ironclad evaluation covers 11 tasks, and the report does not provide enough detail to determine how representative those tasks are of routine contract work. The direction of the reported findings is suggestive, but their scope and causes remain unsettled.

Amazon

software review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch Review Practices and Training

The next useful evidence will be whether organisations publish more detail on how AI-generated work is reviewed, including review rates, time to approval, error rates and the criteria used to judge quality. In mathematics, readers will need to distinguish formally verified results from work that has not received that check, and follow any later corrections or independent assessments. The source does not identify a scheduled review milestone for the full set of manuscripts.

For software and professional services, the key test is whether higher production volume can be matched with sufficient qualified review rather than simply faster approval. Employers will also need to show how junior workers can gain practical experience as AI systems take on more drafting and coding. No common standard or timetable for addressing these issues is identified in the source material.

Amazon

AI verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What happened in mathematics?

OpenAI published 722 mathematical manuscripts from a programme that explored about 4,000 problems. The manuscripts were grouped into 372 families, and the source says the average result took about three hours of compute to produce.

Were all the manuscripts formally verified?

No. The source says some results were formally checked in Lean and quotes OpenAI warning that unformalized results “could have issues.” It does not state how many manuscripts received formal verification.

What does the software data say about review?

The cited analyses report longer review waits and, in some cases, pull requests merged without human review. For example, a peer-reviewed 2026 study cited in the source found 61% of AI-agent pull requests received no human review before merging or closing. The studies use different methods and populations.

Does this prove AI work is less reliable?

No single conclusion follows from these figures. Lower acceptance rates or longer waits may reflect quality concerns, review policies, task differences or other factors. The source does not establish that all AI-generated work is less reliable than human-produced work.

What remains to be seen?

It remains unclear how much AI-generated work receives meaningful review, how often review catches important errors, and whether junior professionals will still get enough hands-on experience to become expert reviewers. More consistent, transparent measures would help answer those questions.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI In Your Car: Anthropic’s Claude Launches On CarPlay Platform

Anthropic’s Claude AI assistant is now available on Apple CarPlay, enabling voice interactions in vehicles and expanding AI presence on dashboards.

Nitter And XCancel Receive Cease And Desist Notices

Nitter and XCancel, privacy-focused Twitter alternatives, receive official cease and desist notices, raising questions about their future and legal challenges.

China’s DeepSeek Debuts V4 Pro, Challenging Global AI Leaders Like Claude Fable 5

DeepSeek has reportedly launched V4 Pro, claiming performance comparable to Anthropic’s Claude Fable 5, but independent verification is not yet available.

A New Way To Find Radisson Hotels: Discovery In ChatGPT

An OpenAI announcement says Radisson hotels are coming to ChatGPT discovery, but launch timing, availability and booking features remain unspecified.