Did GPT-5 really solve previously unsolved math problems?
OpenAI claimed GPT-5 solved hard Erdős problems open for decades. But what did the model actually do, and how was the claim verified or challenged by domain experts?
OpenAI researchers posted on X that GPT-5 had "found solutions to 10 (!) previously unsolved Erdős problems" and made progress on eleven more, problems "open for decades." The first post, by Kevin Weil, was later deleted and others echoed it. Matthias Bastian, reporting for The Decoder on 2025-10-18, says the wording suggested GPT-5 had independently produced proofs of hard number theory questions. Thomas Bloom, who runs erdosproblems.com, called this "a dramatic misinterpretation": on his site "open" means only that he personally does not know the solution. GPT-5 had surfaced existing research he had missed. The original tweets were mostly deleted and the researchers admitted the mistake.
The mechanism is a shift in what "open" refers to. On Bloom's site the word reports the state of one maintainer's knowledge; in the announcement it read as the state of mathematics. Bastian adds that Bubeck "knew what GPT-5 actually contributed, but still used the ambiguous phrase 'found solutions'," so the failure was in the wording of the claim, not in the literature retrieval itself. The excerpt's own account of the useful contribution is retrieval: GPT-5 "proved useful as a research tool for tracking down relevant academic papers," especially where literature is scattered or terminology is inconsistent. Bastian reports that Terence Tao sees the nearest-term value in "tedious tasks like literature searches" rather than the toughest open problems, with only "isolated examples of progress" on hard questions, and that human expertise is still needed to review, classify and integrate AI-generated results.
Against the neighbors, the episode is the verification gap in miniature. A claim of a new result was settled by the person who maintains the record of which problems are open, and that check, not the announcer's confidence, decided what was true. Can a higher evaluation score hide poor task performance? describes the same split between apparent and actual progress, though in an optimization loop rather than a public announcement. How prone is autonomous AI research to reward hacking? treats the validity of AI-produced research as a premise to test; this case shows one way that test gets run in practice, by an outside expert, not by the lab's own review.
The excerpt does not establish much. It is a secondary news report, the original posts are deleted and not quoted in full, and it gives no count of how many of the listed problems were new to the literature, nor any measurement of GPT-5's retrieval. The "10" and "eleven more" are Weil's figures, not checked here. Tao's remarks are relayed through Bastian. What follows, at the strength of one episode, is narrow: a claim about AI progress on research problems should be read as a claim about the status of a problem in a maintained record, and whether it holds depends on a check by someone who knows that record. The episode does not show that AI cannot contribute to open problems.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can we trust AI-generated mathematical proofs without understanding them?- Why do some Erdős problem solutions fail to resolve the originally intended claims?
- Why did the Jacobian conjecture resist proof for over a century?
- What pattern does this follow from OpenAI's earlier Erdős problem claim?
- Why did OpenAI's Erdős primality claim collapse under independent verification?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How prone is autonomous AI research to reward hacking?
When AI agents autonomously optimize research metrics with broad permissions and fuzzy objectives, do they exploit shortcuts that inflate scores without improving actual performance? Understanding this matters for trusting AI-generated research results.
the validity premise; this case shows an outside expert settling a research claim
-
Can a higher evaluation score hide poor task performance?
When language models are optimized to maximize evaluation scores, does that score improvement always reflect genuine progress on the underlying task, or can systems game the evaluator while task performance stagnates or worsens?
the same gap between apparent and actual progress, in a public announcement instead of an optimization loop
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Leading OpenAI researcher announced a GPT-5 math breakthrough that never happened
- Mathematicians are developing rules for AI use — other fields should follow
- Why the Legendary Erdős Problems Are Falling to AI
- The crisis of AI-generated mathematics
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Remarks on the disproof of the unit distance conjecture
- AI companies must work with the research community to protect attribution
Original note title
OpenAI researchers' Erdős claim dissolved into a literature search — open meant unknown to Bloom, not unsolved