Complementary remarks from Gary Marcus and Terence Tao on OpenAI's giant math drop
Source: Gary Marcus (with Terence Tao), Marcus on AI · 2026-10-07
The real news here isn’t the result; it’s not what we were told.
- AI once tried to be a science. Now we get stuff like the completely vague report from OpenAI below:
“Same procedure”? “Using an unreleased model”?
This would never pass peer review.
We don’t know what the procedure was.
We know nothing about the architecture. For esxample, were the proofs generated in one shot, and then verified by the symbolic system Lean? Was there an iterative process?)
We know nothing about the failure rate. We know nothing about the training/post training/data augmention.
As a result, we have zero idea of how generalizable the result is outside math.
A lot of the discussion on social media has been reduced to an ignorant cheering section that applauds without knowing what it is applauding or what it might mean— without ever asking basic scientific questions.
The new system could be a legitimate step toward AGI. Or it could just be a clever leveraging of Lean and synthetic data in a verifiable domain with no generality whatsoever.
From the initial report, we can tell almost nothing.
The press release was the goal.
Mathematic AI slop is not a significant break through.
I posted on x this morning --- Bryna Kra specifically objected to dumping hundreds of results onto GitHub rather than integrating them into normal scientific processes, and said “math by tweet and math by press release” was not a healthy way to sustain the ecosystem.
Lines of inquiry this paper opens 3
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does diversity prevent model convergence on superficial patterns? Can we trust AI-generated mathematical proofs without understanding them? What human oversight must AI research systems have?