🔍 Read the full analysis: How AI Turned Review Into The Expensive Part on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
AI systems are producing research, software changes and contract work faster, but the source material describes human review as a growing constraint on using that output. The reported figures point to a widening gap between production and verification, though several software metrics come from companies that sell review tools and should be read with care.
AI-generated work is arriving faster than people can verify it, according to figures spanning mathematics, software development and contract workflows. The source material says OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems, while examples from software teams show review taking longer or being skipped; the broader concern is that checking may become the constraint on how much AI output organisations can safely use.
The mathematical example contrasts rapid production with expert scrutiny. The source says each result took an average of about three hours of compute. An earlier result from the same programme—a counterexample to an old Erdős conjecture—prompted careful verification by five leading mathematicians. OpenAI’s work includes results formally checked in Lean, a proof-assistant system, but the company cautioned that some results not formalized in Lean could have issues. The number of manuscripts does not, by itself, establish how many are correct, important or independently verified.
Software data cited in the material also points to a review burden, though the figures come from industry analyses with different methods. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin. It said those changes were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several data providers sell code-review tools, making attribution and methodology important.
A peer-reviewed study published in 2026, as described in the material, found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros separately reported a 31.3% rise in merges with zero review during high-adoption periods. In contract work, the source cites an OpenAI partnership with contract-software company Ironclad: GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. That figure indicates progress against the evaluation, but also leaves criteria unmet; the source does not provide the task-by-task results or say how performance was measured.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Sets the Pace for AI Work
If AI makes drafting and coding cheaper, organisations may still be limited by the people who can determine whether the result is fit for use. Human review is not just proofreading: reviewers must judge whether a result addresses the right problem, identify omissions, and sometimes accept professional or legal responsibility. A proof checker can verify a proof against a stated theorem, for example, without deciding whether that theorem is the useful one. Tests can confirm software meets the tests without showing that the tests reflect the actual need.
The consequences can run in opposite directions. With too little review capacity, teams may rubber-stamp or ship work without inspection. With too much suspicion toward machine-generated work, useful changes may wait longer than necessary. The source material cites LinearB as finding 38% of reviewers deliberately deprioritise AI-generated changes. These are reported findings, not proof that all teams behave the same way, but they illustrate how uncertainty about quality can slow adoption even when generation speeds up.
The bottleneck also has a workforce dimension. Senior engineers, lawyers, scientists and auditors generally develop judgement through years of doing the underlying work. If junior staff increasingly review AI drafts instead of producing work themselves, the route by which they gain that experience could narrow. The source presents this as a risk, not an established outcome. It matters because future review capacity depends on training people to become reviewers, not simply on buying more generation tools.
As an affiliate, we earn on qualifying purchases.
From Fast Drafts to Expert Checks
The source frames the change as a common pattern across three types of work: mathematical research, software and professional workflows. In mathematics, formal systems such as Lean can check whether a proof follows from its formal statements. That is distinct from deciding whether the statement is meaningful or whether a result changes the field. The material describes this distinction with the phrase “verification abundance, adjudication scarcity”: machines can help test outputs, while expert interpretation remains limited.
In software, a pull request is a proposed change to a codebase. AI tools can generate or modify code, but review still asks whether the change is correct, safe and appropriate for the wider system. The cited measures cover different periods and populations, so they should not be treated as directly comparable or as a single universal estimate. The source itself cautions that some providers sell code-review products.
The contract example makes a similar distinction between task performance and readiness for use. Meeting an average of 55% of evaluation criteria across 11 tasks is a reported benchmark result, not evidence that a contract can be deployed without human oversight. The source does not specify what each criterion covered, how the evaluation was scored or whether the model’s outputs were used in live contracts.
software review automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The figures do not establish a single causal account of AI’s effect on review. The cited software reports examine different samples and measures, and several come from vendors selling tools in this area. The source does not provide the full study methods, time periods or uncertainty ranges for every statistic. The peer-reviewed study’s finding about unreviewed AI-agent pull requests is reported here as presented in the source material, but its sample and definition of review are not supplied.
For mathematics, the source does not say how many of the 722 manuscripts have been independently checked, how many are formally verified or what proportion contain errors. It also does not give enough detail to evaluate the significance of the reported Ironclad evaluation. It remains unclear how quickly review practices, formal verification and workforce training will adapt, or whether productivity gains will outweigh the added review burden in particular organisations.
AI verification tools for research
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tracking Verification Capacity
The next useful evidence will be independently documented results showing not only how much AI-generated work is produced, but how much is accepted after review, how long checks take and what defects are found. In mathematics, that means clearer accounting of formal verification and expert assessment. In software, comparable studies across organisations could help separate changes caused by AI adoption from other shifts in team practices.
Organisations adopting these systems will also need to decide who has authority to approve outputs and how junior staff can gain the experience needed for that responsibility. The immediate question is not only whether AI can generate a draft, proof or code change, but who checks it, by what standard and who is accountable when it is used. The source material identifies that as the emerging pressure point; it does not establish that any one profession or solution will resolve it.
human review software for AI output
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source describes AI producing research, software changes and contract work at growing scale, while human verification remains comparatively slow and limited. It argues that review is becoming a constraint on practical use, rather than claiming that all AI output is unreliable.
What does the source say OpenAI published?
It says OpenAI published 722 mathematical manuscripts this week after posing about 4,000 problems. The material also says some results were formally checked in Lean and that OpenAI warned unformalized results could have issues.
Do the software figures prove AI code is worse?
No. The figures cited report differences in review waits, acceptance and review practices, but they do not by themselves prove why those differences occurred or apply to every software team. The source also cautions that some data providers sell code-review products.
Why can’t AI simply review AI-generated work?
Automated checks can test outputs against formal rules or specified tests, but they may not determine whether the original requirements were correct or complete. The source also points to human accountability in areas such as contracts and engineering, where a person or institution must stand behind a decision.
What remains unknown?
The source does not establish how many mathematical manuscripts have been independently verified, how the contract evaluation was scored, or whether review capacity will keep pace with AI output. The cited industry figures also use different methods and populations.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
