What procedural details did OpenAI withhold from its math announcement?
Gary Marcus examines what information OpenAI's math result report omitted—method, architecture, failure rates—and whether the gap prevents independent evaluation of the claim's validity and generalizability.
Gary Marcus argues that OpenAI's report of its "giant math drop" omitted the information a reader would need to evaluate it. He quotes the report's own vague phrasing back at it — "Same procedure"? "Using an unreleased model"? — and says flatly "this would never pass peer review." His list of what is missing is specific: "We don't know what the procedure was," "We know nothing about the architecture," including whether proofs were generated in one shot or through an iterative process checked by the Lean symbolic system, and "We know nothing about the failure rate. We know nothing about the training/post training/data augmentation." From that gap he draws a scope claim: "we have zero idea of how generalizable the result is outside math."
Marcus's reasoning treats the announcement as a stand-in for a paper and measures it against what a paper would owe a reader: method, architecture, error rate, and training details, so that a result can be checked and its limits known. Without those, he argues, the claim cannot be placed anywhere on the range between "a legitimate step toward AGI" and "a clever leveraging of Lean and synthetic data in a verifiable domain with no generality whatsoever" — "from the initial report, we can tell almost nothing." He extends the complaint to the discourse around the release, calling social-media reaction "an ignorant cheering section that applauds without knowing what it is applauding," and relays a parallel objection from Bryna Kra, who told him on X that dumping hundreds of results onto GitHub rather than routing them through normal scientific channels made "math by tweet and math by press release" unsustainable for the field.
This sits beside Did GPT-5 really solve previously unsolved math problems? as a second case of an OpenAI math claim that outran what had actually been shown, but the mechanism differs: the Erdős claim collapsed under a named check by the one person who held the relevant record, while Marcus's complaint is that the "giant math drop" announcement supplies no procedural detail for anyone to check against at all. Kra's objection adds a different register — not whether this result holds, but whether announcing results by GitHub dump and press release can sustain mathematics as a shared process, which bears on the verification-labor problem in Does AI-generated mathematics break the link between proof and understanding?: if results arrive with no procedure to review, the "boring" verification work that essay says the field undervalues cannot even begin.
The excerpt is Marcus's commentary, not an independent audit of the underlying result — he names no specific errors in the math itself, only the absence of information needed to find any. Terence Tao is credited as a co-source in the piece's title, but the excerpt as captured contains none of his remarks, so nothing here can be attributed to Tao. What the excerpt supports, at the strength a position piece allows, is narrower than "the result is wrong": a report that withholds procedure, architecture and failure rate cannot be evaluated for correctness or generality by outsiders, whatever the result turns out to be worth.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does diversity prevent model convergence on superficial patterns?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Did GPT-5 really solve previously unsolved math problems?
OpenAI claimed GPT-5 solved hard Erdős problems open for decades. But what did the model actually do, and how was the claim verified or challenged by domain experts?
a second OpenAI math claim that outran its evidence, but collapsed under a named check rather than withheld procedure
-
Does AI-generated mathematics break the link between proof and understanding?
Can a mathematically correct proof generated by AI still certify the understanding that a human mathematician gained? This matters because papers have traditionally vouched for both correctness and the thinking process behind them.
Kra's objection to GitHub-dump publication bears on the undervalued verification labor this essay describes
-
How much human input did OpenAI's Navier-Stokes proof actually require?
OpenAI claimed its model produced a Navier-Stokes proof with minimal human help, but Buckmaster's account suggests the actual process involved substantial team effort, testing, and prompting. Did the public framing match what actually happened?
Evidence for: Buckmaster's account of undisclosed human involvement supports Marcus's charge that OpenAI withheld procedural detail
-
Can AI governance models from mathematics work across scientific fields?
Should the Leiden Declaration on AI and mathematics—a set of principles for responsible AI use—serve as a template for other disciplines? Nature argues it should, using OpenAI's undisclosed unit-distance proof as a test case for why disclosure matters.
Evidence for: Nature's report of OpenAI's undisclosed software and training data corroborates Marcus's charge of withheld procedural detail
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Complementary remarks from Gary Marcus and Terence Tao on OpenAI's giant math drop
- Mathematicians are developing rules for AI use — other fields should follow
- AI companies must work with the research community to protect attribution
- Leading OpenAI researcher announced a GPT-5 math breakthrough that never happened
- Statement on blowup results with Levent Alpöge
- The crisis of AI-generated mathematics
- Leiden Declaration on Artificial Intelligence and Mathematics
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
Original note title
Marcus argues OpenAI's math result announcement withheld the procedural detail scientific review requires