How AI Turned Review Into The Expensive Part
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How AI Turned Review Into The Expensive Part on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

AI systems are producing research, software changes and contract work faster, but the source material describes human review as a growing constraint on using that output. The reported figures point to a widening gap between production and verification, though several software metrics come from companies that sell review tools and should be read with care.

AI-generated work is arriving faster than people can verify it, according to figures spanning mathematics, software development and contract workflows. The source material says OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems, while examples from software teams show review taking longer or being skipped; the broader concern is that checking may become the constraint on how much AI output organisations can safely use.

The mathematical example contrasts rapid production with expert scrutiny. The source says each result took an average of about three hours of compute. An earlier result from the same programme—a counterexample to an old Erdős conjecture—prompted careful verification by five leading mathematicians. OpenAI’s work includes results formally checked in Lean, a proof-assistant system, but the company cautioned that some results not formalized in Lean could have issues. The number of manuscripts does not, by itself, establish how many are correct, important or independently verified.

Software data cited in the material also points to a review burden, though the figures come from industry analyses with different methods. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin. It said those changes were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several data providers sell code-review tools, making attribution and methodology important.

A peer-reviewed study published in 2026, as described in the material, found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros separately reported a 31.3% rise in merges with zero review during high-adoption periods. In contract work, the source cites an OpenAI partnership with contract-software company Ironclad: GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. That figure indicates progress against the evaluation, but also leaves criteria unmet; the source does not provide the task-by-task results or say how performance was measured.

At a glance
analysisWhen: Developing; the source material cites O…
The developmentA report on AI’s expanding output argues that the scarce resource is shifting from producing work to verifying and taking responsibility for it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Sets the Pace for AI Work

If AI makes drafting and coding cheaper, organisations may still be limited by the people who can determine whether the result is fit for use. Human review is not just proofreading: reviewers must judge whether a result addresses the right problem, identify omissions, and sometimes accept professional or legal responsibility. A proof checker can verify a proof against a stated theorem, for example, without deciding whether that theorem is the useful one. Tests can confirm software meets the tests without showing that the tests reflect the actual need.

The consequences can run in opposite directions. With too little review capacity, teams may rubber-stamp or ship work without inspection. With too much suspicion toward machine-generated work, useful changes may wait longer than necessary. The source material cites LinearB as finding 38% of reviewers deliberately deprioritise AI-generated changes. These are reported findings, not proof that all teams behave the same way, but they illustrate how uncertainty about quality can slow adoption even when generation speeds up.

The bottleneck also has a workforce dimension. Senior engineers, lawyers, scientists and auditors generally develop judgement through years of doing the underlying work. If junior staff increasingly review AI drafts instead of producing work themselves, the route by which they gain that experience could narrow. The source presents this as a risk, not an established outcome. It matters because future review capacity depends on training people to become reviewers, not simply on buying more generation tools.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Fast Drafts to Expert Checks

The source frames the change as a common pattern across three types of work: mathematical research, software and professional workflows. In mathematics, formal systems such as Lean can check whether a proof follows from its formal statements. That is distinct from deciding whether the statement is meaningful or whether a result changes the field. The material describes this distinction with the phrase “verification abundance, adjudication scarcity”: machines can help test outputs, while expert interpretation remains limited.

In software, a pull request is a proposed change to a codebase. AI tools can generate or modify code, but review still asks whether the change is correct, safe and appropriate for the wider system. The cited measures cover different periods and populations, so they should not be treated as directly comparable or as a single universal estimate. The source itself cautions that some providers sell code-review products.

The contract example makes a similar distinction between task performance and readiness for use. Meeting an average of 55% of evaluation criteria across 11 tasks is a reported benchmark result, not evidence that a contract can be deployed without human oversight. The source does not specify what each criterion covered, how the evaluation was scored or whether the model’s outputs were used in live contracts.

Amazon

software review automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The figures do not establish a single causal account of AI’s effect on review. The cited software reports examine different samples and measures, and several come from vendors selling tools in this area. The source does not provide the full study methods, time periods or uncertainty ranges for every statistic. The peer-reviewed study’s finding about unreviewed AI-agent pull requests is reported here as presented in the source material, but its sample and definition of review are not supplied.

For mathematics, the source does not say how many of the 722 manuscripts have been independently checked, how many are formally verified or what proportion contain errors. It also does not give enough detail to evaluate the significance of the reported Ironclad evaluation. It remains unclear how quickly review practices, formal verification and workforce training will adapt, or whether productivity gains will outweigh the added review burden in particular organisations.

Amazon

AI verification tools for research

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Verification Capacity

The next useful evidence will be independently documented results showing not only how much AI-generated work is produced, but how much is accepted after review, how long checks take and what defects are found. In mathematics, that means clearer accounting of formal verification and expert assessment. In software, comparable studies across organisations could help separate changes caused by AI adoption from other shifts in team practices.

Organisations adopting these systems will also need to decide who has authority to approve outputs and how junior staff can gain the experience needed for that responsibility. The immediate question is not only whether AI can generate a draft, proof or code change, but who checks it, by what standard and who is accountable when it is used. The source material identifies that as the emerging pressure point; it does not establish that any one profession or solution will resolve it.

Amazon

human review software for AI output

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The source describes AI producing research, software changes and contract work at growing scale, while human verification remains comparatively slow and limited. It argues that review is becoming a constraint on practical use, rather than claiming that all AI output is unreliable.

What does the source say OpenAI published?

It says OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems. The material also says some results were formally checked in Lean and that OpenAI warned unformalized results could have issues.

Do the software figures prove AI code is worse?

No. The figures cited report differences in review waits, acceptance and review practices, but they do not by themselves prove why those differences occurred or apply to every software team. The source also cautions that some data providers sell code-review products.

Why can’t AI simply review AI-generated work?

Automated checks can test outputs against formal rules or specified tests, but they may not determine whether the original requirements were correct or complete. The source also points to human accountability in areas such as contracts and engineering, where a person or institution must stand behind a decision.

What remains unknown?

The source does not establish how many mathematical manuscripts have been independently verified, how the contract evaluation was scored, or whether review capacity will keep pace with AI output. The cited industry figures also use different methods and populations.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple’s new SpeechAnalyzer API, benchmarked against Whisper and its predecessor

Apple’s new SpeechAnalyzer API is tested against Whisper and its predecessor, highlighting its performance and potential applications.

Baidu’s Unlimited-OCR: Separating Fact From Fiction In AI Tech

Baidu released Unlimited-OCR, a 3-billion-parameter model capable of parsing multi-page documents in a single pass. This report clarifies what is confirmed and what remains uncertain.

Fable 5 Is Back. GPT-5.6 Is Next. And Anthropic Reportedly Already Has Something Stronger.

Anthropic restores Fable 5 after government blackout; OpenAI previews GPT-5.6, and rumors suggest an even more advanced model exists privately.

The Agent Trap: Why 90% of AI “Launches” Are Infrastructure Liars

Most AI ‘agent’ launches in 2026 are features on existing infrastructure, not true autonomous agents. This shift impacts enterprise AI procurement and security.