SYNTHESIS NOTE
Topics›Correct but Not Understood›this note

What does correctness of outputs tell us about reasoning?

IMO graders verified that Gemini's proofs were mathematically correct, but their review excluded the model's processes and training. This raises whether certified right answers demonstrate genuine understanding or only output accuracy.

Synthesis note · 2026-10-06 · sourced from Correct but Not Understood

Google DeepMind claims that an advanced version of Gemini Deep Think solved five of the six problems at the 2025 International Mathematical Olympiad, earning 35 of 42 points, a gold-medal score. The result is the lab's own claim, but the excerpt says the answers were graded by IMO coordinators "using the same criteria as for student solutions," and that the IMO "confirmed that our submitted answers are complete and correct solutions." That is the verified part: the proofs were checked by the competition's own graders. The only comprehension-like statement is that the graders found the solutions "clear, precise and most of them easy to follow." That describes readability. The excerpt does not say the reasoning was explained or understood in human terms.

The excerpt credits the result to a reasoning mode that "simultaneously explore[s] and combine[s] multiple possible solutions before giving a final answer, rather than pursuing a single, linear chain of thought." The model was also trained with novel reinforcement learning on multi-step reasoning and theorem-proving data, given a curated corpus of solutions, and given general hints in its instructions. The excerpt contrasts this with 2024, when AlphaGeometry and AlphaProof needed problems translated into Lean and two to three days of computation. This year the model worked end-to-end in natural language within the 4.5-hour limit.

The parallel-exploration mechanism is a different axis from the result in Does more thinking time always improve reasoning accuracy?, which concerns what happens when a single chain is made longer. The excerpt attributes its gains to breadth of exploration, not length, and gives no measurement that would test the non-monotonic finding. The certification also bears on Do automated benchmarks hide what frontier AI systems can really do?. An expert-graded competition is a stronger check than an automatically graded benchmark, but it still stops at the output. The accuracy of a result says little about how it was reached, which is the gap Can we measure how deeply a model actually reasons? tries to close from inside the model.

The excerpt does not establish how the system reached its proofs, and the lab says so itself: the IMO's "review does not extend to validating our system, processes, or underlying model." Five correct proofs on one competition's problems show that the outputs were correct on that occasion. They do not show that the training data, the hints or the parallel procedure produced the capability, or that it transfers beyond olympiad problems. At the strength the evidence allows, "correct" holds for these proofs, while "understood" does not yet hold for the system that produced them, and the public claim should be read at that strength.

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can we trust AI-generated mathematical proofs without understanding them? How do educators verify student capability when AI can produce indistinguishable work?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 135 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Google DeepMind reports IMO graders certified the Gemini Deep Think proofs as correct without validating the system behind them