How much human input did OpenAI's Navier-Stokes proof actually require?
OpenAI claimed its model produced a Navier-Stokes proof with minimal human help, but Buckmaster's account suggests the actual process involved substantial team effort, testing, and prompting. Did the public framing match what actually happened?
Tristan Buckmaster's public statement describes a dispute over a rumored OpenAI result on the Navier-Stokes blowup problem, timed against his and Levent Alpöge's own LLM-assisted blowup results for porous media, Boussinesq, and Euler. Sebastien Bubeck told Buckmaster that an internal OpenAI model had produced a proof of finite-time blowup for forced Navier-Stokes, and that Levent "had been told by Sebastien 'very little human input' had been used." Buckmaster writes plainly: "This turned out not to be true."
Buckmaster gives the specifics that undercut the "very little human input" framing: as members of the OpenAI team supplied Bubeck with corrections over an internal chat during the call, it emerged that "an entire team had been working on the problem," that the team had first tested the model on easier problems including Euler, that "even the prompt that had been shown to me had been written by prompting Codex," and that the first prompt was sent only "in the past few days, after information about our work had reached OpenAI" — a fact Buckmaster says OpenAI did not answer directly for some time. He also reports pressure that went beyond the mathematics: a proposal to remove Alpöge from authorship because "it was so annoying" that he works at Anthropic, and a reply of "Why would you ruin your career?" when Buckmaster said he would go public.
This sits beside Did GPT-5 really solve previously unsolved math problems? as a second instance of an OpenAI math capability claim needing correction once more of the story came out — but the gap this time is not that a problem was already known to specialists, it is that the degree of human scaffolding behind a model's output was represented as smaller than it was. It also complicates Can opaque machine learning models help prove new mathematics?: Tao's framing assumes validation of a model's output is the open question, whereas Buckmaster's account shows the validation he was denied was of the claim about the process — how much the model was told, tested, and prompted by people — not just the proof itself.
Buckmaster is explicit about the limits of his own claim: "I have not seen OpenAI's proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me." The excerpt therefore establishes a credible first-person account of a contested, commercially-pressured framing of human involvement, not a verified fact about what the OpenAI model actually did unaided. The implication the evidence does support is narrower but still real: when a lab's own account of "minimal human input" is the only source for a capability claim, and that account changes once the people who built the result start correcting it on a call, the claim was not fact-checkable from the outside at the time it was first represented.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can we trust AI-generated mathematical proofs without understanding them? How does diversity prevent model convergence on superficial patterns? What human oversight must AI research systems have?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Did GPT-5 really solve previously unsolved math problems?
OpenAI claimed GPT-5 solved hard Erdős problems open for decades. But what did the model actually do, and how was the claim verified or challenged by domain experts?
an earlier OpenAI math capability claim that also shrank once more of the context surfaced
-
Can opaque machine learning models help prove new mathematics?
Tao explores whether ML tools' opacity disqualifies them from research mathematics, and under what conditions their suggestions might be trustworthy enough to guide rigorous proofs.
Buckmaster's account shows the contested validation was of the human-input claim, not only the proof
-
Does Sakana's AI Scientist deliver autonomous research without human help?
Can an AI system truly run the complete research lifecycle alone, or does it still need human guidance and oversight? This matters for understanding whether automated research can scale.
another case of an autonomous-AI-research claim overstating what the system did on its own
-
Can we trace AI contributions to scientific breakthroughs?
When AI systems help produce major research results, how can we identify what training data or prior work actually contributed? The Buckmaster-OpenAI dispute shows current systems have no way to track this.
Extends the dispute into policy: Nature cites it to urge opt-in data sharing and attribution safeguards from AI firms
-
What procedural details did OpenAI withhold from its math announcement?
Gary Marcus examines what information OpenAI's math result report omitted—method, architecture, failure rates—and whether the gap prevents independent evaluation of the claim's validity and generalizability.
Extends the same dispute: Marcus says OpenAI's announcement omitted procedure, architecture, and failure-rate details needed for review
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AI companies must work with the research community to protect attribution
- Statement on blowup results with Levent Alpöge
- Culture Becomes a Dark Forest
- A Bad Day For Humans, a Worse Day For Humanity
- Complementary remarks from Gary Marcus and Terence Tao on OpenAI's giant math drop
- The crisis of AI-generated mathematics
- Mathematicians are developing rules for AI use — other fields should follow
- Machine-Assisted Proof
Original note title
Buckmaster says OpenAI's claim that its Navier-Stokes proof used very little human input did not hold up