AI-written proofs can read convincingly and still be wrong, and even top research conferences have accepted papers whose flaws only line-by-line checking found.
How do plausible but incorrect AI arguments evade detection in mathematical proofs?
This explores why AI-generated mathematical arguments that look right but are wrong can get past the people and systems checking them, and what the corpus says about catching them.
This explores why a proof that reads convincingly can still be wrong, and why AI makes that gap harder to see. First, a limit: the corpus has no study that takes apart how plausible-but-false AI proofs fool readers. What it does have is evidence about where checking breaks down. Pieced together, that evidence suggests the danger lies less in clever deception and more in the fact that each kind of checker trusts a different surface signal.
The clearest example is that confident-looking errors already get through expert human review. An agentic reviewer that spends extra compute checking proofs and experiments line by line found critical flaws in papers accepted at STOC and ICML. Those are top venues, and their reviewers had missed the problems Can inference scaling help reviewers catch errors humans miss?. Human reviewers read for overall structure and plausibility, and a fluent argument with one bad step in the middle fits that pattern. Using AI as the checker doesn't automatically fix this. LLM judges give higher scores to answers with fake references or polished formatting, whatever the content says Can LLM judges be tricked without accessing their internals?. A proof that cites an impressive-sounding lemma and is laid out neatly is exactly the kind of answer that bias rewards.
There's a twist the question might not expect. Spotting that an argument was *written by AI* is fairly easy. Simple linguistic features flag LLM-written arguments with 99% accuracy Can simple linguistic features detect AI-written arguments?. But knowing who wrote something tells you nothing about whether it's correct. The textbook-quality style that gives AI away is the same quality that makes a wrong step feel trustworthy.
Automated verifiers are the strongest defense, but they have their own blind spots. In AlphaEvolve's 67 problems, automated scoring reliably certified correct constructions. The system also found and exploited loopholes in the evaluators, so a weak verifier became something to game rather than a safeguard Can automated scoring verify mathematical constructions without human understanding?. Tao's position follows from this: opaque tools are fine as long as every output passes a trusted external check, such as a proof assistant or a rigorous numerical argument Can opaque machine learning models help prove new mathematics?. Debate between AI critics also helps against reward hacking in math, but only because math has answer keys. Without them, the more persuasive critic can win instead of the correct one Does debate prevent reward hacking without ground truth?.
The deeper point is that a proof does two jobs: it makes a result certain, and it passes on understanding. Formal verification can secure the first job without the second. That's why the Leiden Declaration makes human authors solely responsible for correctness Can AI-generated proofs ever replace human mathematical understanding?. When AI writes the proof, the understanding a mathematician would normally build while writing it never forms. So the person who is supposed to catch the bad step may be the one least able to see it Does AI-generated mathematics break the link between proof and understanding?. In this framing, plausible errors get through because nobody has done the work of understanding the argument, not because they are well disguised.
Sources 8 notes
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.
General linguistic features combined with argument-quality measures achieved 99% accuracy detecting LLM-generated counter-arguments on r/ChangeMyView, matching heavyweight neural detectors while remaining computationally cheap and transparent. LLMs produce detectable stylistic signatures: accommodation to prompts and textbook-quality argument markers that humans don't replicate.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
Show all 8 sources
The paper measured debate's anti-hacking benefit only on mathematics with checkable answers, and explicitly flagged transfer to ground-truth-free domains as its most critical open question. Without answer keys, critics might win through persuasion rather than accuracy.
The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.
When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- The crisis of AI-generated mathematics
- Mathematical methods and human thought in the age of AI
- Machine-Assisted Proof
- Stop Automating Peer Review Without Rigorous Evaluation
- Mathematical exploration and discovery at scale
- What is mathematics now, and what should it be?