AI At Scale: More Creation, More Need For Checking
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI At Scale: More Creation, More Need For Checking on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article describes a growing gap between AI-generated work and the capacity to check it, citing examples in mathematics, software and contract workflows. The figures point to review pressure, though some software data comes from vendors that sell code-review tools, and the source does not provide independent confirmation for every claim.

A report published this week argues that AI is making work cheaper to produce without making it equally cheap to verify, a widening gap illustrated by new mathematical manuscripts and software review data. The examples suggest that human checking capacity may constrain how much AI-generated work organisations can safely use, though several figures come from companies with commercial interests in code review.

The report says OpenAI published 722 mathematical manuscripts this week, selected from work on roughly 4,000 problems. The manuscripts fall into 372 families, and the average result reportedly took about three hours of compute to produce. Some results have been checked in the Lean proof assistant; OpenAI cautioned that unformalized results could have issues. The source does not provide a full account of how each manuscript was assessed before publication.

It contrasts that output with the response to an earlier result from the same programme: a counterexample to an old Erdős conjecture that, according to the report, was carefully verified by five leading mathematicians. That example captures the report’s central point: producing candidate results can scale rapidly, while deciding whether a result is correct, relevant and meaningful still requires specialist scrutiny.

Software figures in the report point to similar pressure. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found 61% of AI-agent pull requests received no human review before being merged or closed. These studies measure different things, and the report notes that some data providers sell code-review tools.

A separate example comes from contract work. The report says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. Across 11 tasks, the model met an average of 55% of evaluation criteria, which the source describes as a marked improvement over its predecessor. The score also indicates that substantial criteria remained unmet; the report does not specify how the evaluation was designed or whether those tasks represent typical legal work.

At a glance
reportWhen: Report published this week; cited softw…
The developmentA report argues that AI is increasing the volume of work faster than organisations can verify it, using mathematical manuscripts, software pull requests and contract tasks as examples.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Sets the Pace

If AI increases output but review capacity does not keep pace, the practical limit on adoption may shift from generating work to judging whether that work is usable. Organisations can receive more code, research and draft contracts, but each item still needs an appropriate level of checking before it can be relied on. That can create delays, skipped reviews or extra work for experienced staff.

The report also raises a workforce concern: reviewers often gain expertise by doing the work they later evaluate. If junior employees mainly supervise machine-produced drafts rather than learning through drafting, coding or proving results themselves, the future supply of experienced reviewers could weaken. That is a concern raised by the source, not a demonstrated outcome of the cited data.

In that scenario, specialist judgement becomes a constraint on the value organisations can extract from AI. Senior engineers, lawyers, auditors and researchers may face more review work even as some production tasks become faster. The cited evidence does not establish how pay or employment will change, but it highlights a question for employers: whether they are investing in the people and processes needed to check increased output.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Review Gap

The source frames the examples as a pattern across mathematics, software and professional workflows, rather than as one isolated product announcement. In mathematics, formal tools can check whether a proof follows from its stated assumptions. In software, tests can show whether code meets defined conditions. In contract work, an evaluation can compare a draft against specified criteria. Each method checks something measurable, but none automatically establishes that the original task, assumptions or criteria were the right ones.

The report calls this distinction verification versus adjudication. Verification can establish whether an output meets a formal test; adjudication requires deciding whether the test captures the real need, whether the result matters and who is responsible for approving its use. The source says that an earlier mathematical counterexample became disputed in part because of a gap between proving a stated claim and evaluating the claim itself. It does not provide enough detail here to independently assess that episode.

Across the cited software findings, the measures differ: pull-request volume, time until review begins, acceptance rates and whether human review occurred. They should not be treated as interchangeable or as proof that AI-written code is inherently unreliable. Taken together, the findings reported by the source indicate that more AI-assisted submissions can place pressure on review workflows, while the size and causes of that pressure require careful interpretation.

“Code got cheaper to write and more expensive to trust.”

— ThorstenMeyerAI.com source article

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The source does not provide the complete methods, definitions or underlying datasets for every figure, so comparisons across organisations and studies remain limited. In particular, the cited software data includes providers that sell review products, a potential commercial interest readers should keep in mind. The report also does not establish whether longer review waits result from AI use itself, changes in workload, team practices or other factors.

For the mathematics and contract examples, important details are also missing: how all 722 manuscripts were selected and reviewed, which results were formally verified, and how the 11 contract tasks and their evaluation criteria were chosen. The cited examples illustrate possible limits of automated checking, but do not measure the total cost or accuracy of human review against AI output across these fields.

It is also not yet clear whether organisations will respond by expanding review teams, improving automated checks or accepting more unreviewed work. The source warns about a possible decline in apprenticeship opportunities, but provides no longitudinal evidence that this has already reduced the supply of experienced reviewers.

Amazon

contract review automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The immediate test for the report’s argument is whether further, independently documented evidence shows review time rising as AI-assisted output increases. In software, that means following not just submission volumes but also review coverage, defects, rework and the reasons changes are delayed or accepted. For mathematics and contracting, clearer information about verification methods and evaluation criteria would help readers judge what the reported results establish.

Employers will also need to decide how junior staff acquire the expertise required for oversight. Possible approaches include assigning trainees substantive drafting and coding work, pairing them with senior reviewers, and recording who approved consequential outputs. These are potential responses, not practices the source says have been adopted across the industries it discusses.

For now, the central development is a growing set of examples in which AI-generated work is accumulating faster than human review can readily absorb it. The scale of that gap, and whether new checking tools can narrow it without weakening accountability or training, remain open questions.

Amazon

AI project management tools for verification

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the report’s main finding?

It argues that AI can increase the volume of produced work faster than people can verify it. The examples cover mathematical manuscripts, software pull requests and contract tasks; they do not establish the same effect in every field.

Did OpenAI publish 722 mathematically verified results?

The source says OpenAI published 722 manuscripts drawn from work on about 4,000 problems. It says some results were formally checked in Lean and quotes OpenAI warning that some unformalized results could have issues; it does not say all 722 were formally verified.

What do the software figures show?

The cited studies report increased pull-request volume alongside longer review waits, lower acceptance rates for AI-generated changes in one dataset, and many AI-agent pull requests receiving no human review in another. The studies use different measures, and some data providers sell code-review tools.

Does the report prove AI output is unreliable?

No. The examples raise questions about review capacity and checking, but the source does not provide a single independent comparison that proves AI output is generally unreliable. The specific quality of work depends on the task, evaluation method and review process.

What remains unknown?

The full methods behind several reported figures, how representative the examples are, and whether increased review pressure is reducing training opportunities are not established in the source. It is also unclear how much automated checking can reduce the need for accountable human judgement.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

9 Best Usb Microphones That Incorporate AI For Superior Audio In 2026

Discover the nine best USB microphones with AI features in 2026, offering superior sound quality for streaming, podcasting, and professional use.

Upgrade Your Notes With 7 Top AI Apps In 2026

Discover the 7 best AI-powered note apps in 2026, featuring advanced transcription, summarization, and device compatibility to boost your note-taking.

Applied Materials, Teradyne, and Entegris Stocks Trade Down, What You Need To Know

Shares of Applied Materials, Teradyne, and Entegris decline on broader market worries and sector-specific pressures, raising questions for investors.