INQUIRING LINE

Many claimed AI math wins turn out to be known results rediscovered, and a right answer can't show which happened.

What distinguishes rediscovering known results from genuine mathematical research?

This explores what separates an AI system that re-derives mathematics already known from one that actually advances it, and why the difference is harder to see than it sounds.


This explores what separates an AI that re-derives known mathematics from one doing real research, and why a correct answer can't tell you which one happened. The bluntest finding in the corpus is that many claimed theorem-proving successes turn out to be rediscoveries of existing results. Current systems work as solvers of isolated, well-defined problems rather than as research agents that can take on open questions Can LLM theorem provers tackle genuinely open-ended research problems?. The same note adds a quieter warning: a formally verified proof can still prove something slightly different from the claim mathematicians actually cared about. Verification checks the logic. It doesn't check whether you asked the right question.

Rediscovery is hard to spot partly because 'solving' can be remembering. One model could rebuild 54.6% of a popular math benchmark from partial prompts, then scored 0% on problems released after its training Does RLVR success on math benchmarks reflect genuine reasoning improvement?. Other work found that math performance falls apart when only the numbers in a problem change, or when an irrelevant sentence is added Does LLM math reasoning truly generalize or just pattern match?. Even training that clearly improves reasoning mostly makes each step follow cleanly from the last. A chain of steps can read smoothly and still be wrong as a whole proof Does RLVR actually improve mathematical reasoning or just coherence?. So a fluent, correct-looking derivation is weak evidence that anything new happened. Gemini's IMO result shows the limit clearly: the graders certified that the proofs were correct and explicitly declined to say anything about how the system got there What does correctness of outputs tell us about reasoning?.

What does real research look like, then? Terence Tao's answer shifts the focus from the model to the pipeline around it. An opaque neural network can propose something new, like candidate solutions for blowup in fluid equations, as long as an independent check (a proof assistant, a numerical method, a perturbation argument) confirms what it suggests Can opaque machine learning models help prove new mathematics?. On this view, novelty comes from the system's ability to reach places nobody has been, plus a trustworthy way to confirm it got there. It's suggestive that LLM-generated research ideas get rated as more novel than experts' ideas but less feasible Do language models generate more novel research ideas than experts?. Reaching for new ground is cheap. Landing there with something sound is the hard part.

The less obvious answer is that mathematicians may not define research by its outputs at all. The Leiden Declaration holds that a proof does two jobs: it establishes certainty, and it conveys understanding. Formal verification can only secure the first, so the declaration places credit and responsibility with human authors Can AI-generated proofs ever replace human mathematical understanding?. A related essay argues that AI-generated proofs break the old link between a correct paper and the insight its author gained by writing it Does AI-generated mathematics break the link between proof and understanding?. Much of that insight lives in the dead ends, and papers routinely leave the failed branches out Can research papers preserve the experiments that failed?. That points to a strange conclusion. The finished proof alone may never tell you whether research took place. The evidence is in the process: what was tried, what failed, and what someone came to understand along the way.


Sources 10 notes

Can LLM theorem provers tackle genuinely open-ended research problems?

Current systems excel at isolated, well-defined proofs but cannot address truly open problems like Millennium Prize Problems. Many claimed successes rediscover existing results, and formal verification does not guarantee the proof addresses the intended mathematical claim.

Does RLVR success on math benchmarks reflect genuine reasoning improvement?

Qwen2.5-Math-7B reconstructs 54.6% of MATH-500 from partial prompts but scores 0.0% on post-release LiveMathBench, revealing dataset contamination. On clean benchmarks, only correct rewards improve performance; random and inverse rewards fail or degrade reasoning ability.

Does LLM math reasoning truly generalize or just pattern match?

GSM-Symbolic found that LLMs show high variance across question reformulations, decline sharply when numbers change, and fail when irrelevant but related clauses are inserted. These failures indicate probabilistic pattern-matching rather than true symbolic reasoning.

Does RLVR actually improve mathematical reasoning or just coherence?

RLVR post-training measurably reduces logical errors between adjacent reasoning steps, but locally coherent traces can still be globally invalid proofs. The improvement is structural rather than semantic.

What does correctness of outputs tell us about reasoning?

Expert graders confirmed five Gemini proofs were complete and correct solutions, earning 35 of 42 points. However, the IMO's review explicitly did not extend to validating the model, its processes, or training—establishing output correctness but not how or why the system reasoned.

Show all 10 sources
Can opaque machine learning models help prove new mathematics?

Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

Can AI-generated proofs ever replace human mathematical understanding?

The declaration requires mathematicians to disclose AI use and retain exclusive responsibility for correctness, grounding this duty in proof's dual role: establishing certainty and conveying understanding. Formal verification alone cannot secure both goods.

Does AI-generated mathematics break the link between proof and understanding?

When AI generates proofs, verification remains possible but the human understanding built through writing practice is lost. Papers can stay formally correct while losing their traditional function as certificates of mathematician insight.

Can research papers preserve the experiments that failed?

Publishing imposes a Storytelling Tax (erasing process, failed branches, tacit reasoning) and Engineering Tax (omitting implementation specs). Agent-Native Research Artifacts address both by packaging logic, executable code, exploration graphs of failures, and evidence grounding—treating rejected branches as publishable deliverables rather than editorial casualties.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.