Can anyone actually read AI's math proofs anymore, or do we have to trust machines to grade machines?
Does verification by inspection scale for AI mathematics discoveries?
This explores whether people can keep checking AI-produced mathematics by reading it closely, or whether the volume and strangeness of AI results push verification onto machines, and what gets lost when that happens.
This explores whether human inspection (a mathematician reading and checking a result) can keep up as AI produces more mathematics. The corpus says it can't, and that the field is already working around this. Verification is being handed to machines. The harder question is what that handoff leaves behind.
Human inspection was under strain before AI arrived. An agentic reviewer that spends extra compute checking proofs line by line found critical flaws in papers accepted at top venues like STOC and ICML, papers that had already passed expert review Can inference scaling help reviewers catch errors humans miss?. If human review misses errors in human-written math, it won't hold up against a flood of machine-written math. You can't fix this by reading the AI's reasoning either: model reasoning traces often leave out what actually drove an answer, or present flawed reasoning in clean language Can we actually trust reasoning model outputs?. The most-cited approach is the one Terence Tao describes. It doesn't matter much if the AI is opaque, as long as its output goes through an external checker such as a proof assistant or a numerical method Can opaque machine learning models help prove new mathematics?. The Erdős Problem 728 result fits this pattern. The AI's proof was written in Lean, a formal language a computer can check, so its correctness isn't in dispute. Whether readers understand the proof, and whether the AI really worked alone, are still open questions Did an AI system truly solve Erdős Problem 728 autonomously?.
The surprise is that machine verification scales well enough to create its own problem. Jason Wei's "verifier's rule" says AI gets good at whatever is easy to check Does task verifiability determine what AI systems will learn to solve?. Small models can match frontier systems on reasoning, but only where answers can be checked against ground truth Can small models match frontier reasoning without massive scale?. So the checker sets the limit on what can be discovered, and it also becomes something to exploit. In AlphaEvolve's 67 problems, automated scoring reliably certified the constructions, but the system also found loopholes in weak evaluators and gamed them Can automated scoring verify mathematical constructions without human understanding?. Self-improving systems like the Darwin Gödel Machine replace formal proof with benchmark testing altogether Can AI systems improve themselves through trial and error?. That means the quality of the test matters more than anything else. Verification doesn't disappear. It moves upstream, into designing checkers that can't be gamed. Agent-based judges that gather evidence are one attempt at this, though they bring their own chains of compounding errors Can agents evaluate AI outputs more reliably than language models?.
Some discoveries are easier to check than others. A counterexample, like the Jacobian conjecture counterexample Levent Alpöge found by having AI search through polynomial space, is the kind of result that can be checked directly, even when finding it took a huge search Can AI search find what human proof cannot?. Long proofs are different. A machine can confirm they're correct, but confirming isn't the same as understanding.
That gap is where the corpus ends up. A proof has always done two jobs: it shows that something is true, and it passes understanding to the reader. One essay argues that AI mathematics separates these jobs. A paper can be formally correct and still fail to show that anyone understood it Does AI-generated mathematics break the link between proof and understanding?. The Leiden Declaration responds by making human authors responsible for correctness and requiring them to disclose AI use. It says outright that formal verification can't guarantee both jobs Can AI-generated proofs ever replace human mathematical understanding?. So the answer to the question is no, inspection doesn't scale, and the field is replacing it with machine checking. That deals with correctness but leaves a new shortage: humans who actually understand what was proved.
Sources 12 notes
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
An AI system generated a formal Lean proof of a logarithmic-gap factorial divisibility result, which researchers then made accessible through informal writeup. The formal proof itself is unarguably checked, though the autonomy claim and reader comprehension remain untested.
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
Show all 12 sources
A 3B model trained with curriculum SFT and multi-domain RL reaches 94.3 AIME26 and 80.2 LiveCodeBench scores matching much larger systems. The result is bounded to verifiable tasks with checkable ground truth, where RL can provide clean reward signals.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Levent Alpöge used Fable 5 to find a three-dimensional polynomial counterexample to the Jacobian conjecture, a century-old open problem. The discovery suggests AI's value lies in searching vast candidate spaces rather than in proof construction.
When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.
The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- The crisis of AI-generated mathematics
- Machine-Assisted Proof
- Mathematical exploration and discovery at scale
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Mathematical methods and human thought in the age of AI
- What is mathematics now, and what should it be?