TOPIC

Correct but Not Understood

A subject the collection covers, read through 24 synthesis notes.


View as

Do AI research agents explore as broadly as human researchers?

When AI agents generate research ideas across fields, do they spread their exploration as widely as humans do, or do they cluster around their starting points? This matters for understanding whether AI can drive genuine scientific discovery.

Explore related Read →

Can humans learn chess concepts that AlphaZero discovered alone?

Do grandmasters improve on new puzzles after seeing AlphaZero's solutions to similar positions? The question tests whether superhuman chess knowledge can transfer from machine to human player.

Explore related Read →

How many unit distances can points in a plane have?

The unit distance problem asks how many pairs of points in a set of n points can be exactly distance 1 apart. A recent construction uses number field lattices to achieve an explicit lower bound around n^1.014, improving on a prior result with an unspecified exponent.

Explore related Read →

Can artificial intelligence ever truly understand science?

Researchers ask whether AI can move beyond predicting outcomes to genuinely grasping the theories behind them. The question hinges on what scientific understanding actually means.

Explore related Read →

Can automated scoring verify mathematical constructions without human understanding?

When evolutionary AI systems propose mathematical solutions, does an automated evaluator's score prove correctness sufficiently? The gap between verification and interpretation matters for trust and generalization.

Explore related Read →

Does Sakana's AI Scientist deliver autonomous research without human help?

Can an AI system truly run the complete research lifecycle alone, or does it still need human guidance and oversight? This matters for understanding whether automated research can scale.

Explore related Read →

Did superhuman AI actually improve Go players' decision quality?

After AlphaGo's breakthrough, professional Go players made better moves and tried more novel strategies. But did AI exposure directly cause this improvement, or did players simply memorize AI moves?

Explore related Read →

Can opaque models guide discovery without needing interpretation?

Does deep learning need to be interpretable when it steers hypothesis formation rather than standing as a justified claim itself? The distinction matters for when opacity becomes an epistemic problem.

Explore related Read →

How does FunSearch actually verify its discovered programs?

FunSearch claims to produce verifiable knowledge through evolutionary search paired with an LLM. But what constitutes verification here—scoring functions, human understanding, or formal proof?

Explore related Read →

What does correctness of outputs tell us about reasoning?

IMO graders verified that Gemini's proofs were mathematically correct, but their review excluded the model's processes and training. This raises whether certified right answers demonstrate genuine understanding or only output accuracy.

Explore related Read →

Can predicting scientists improve discovery forecasts?

Does modeling the researchers behind discoveries—their collaborations, expertise, and positioning—forecast future scientific breakthroughs better than analyzing research content alone? This matters because it tests whether human expertise patterns reveal signals that text cannot.

Explore related Read →

How often do language models take specification-violating shortcuts?

ImpossibleBench deliberately creates contradictions between test cases and specs to measure when models cheat rather than solve genuinely. The question explores whether stronger models are more prone to taking shortcuts.

Explore related Read →

Why did Erdős problems become a popular AI testing ground?

Explores what makes Erdős problems attractive for evaluating large language models, including their mathematical domains, difficulty range, and the collaborative infrastructure that enables testing.

Explore related Read →

Can AI-generated proofs ever replace human mathematical understanding?

The Leiden Declaration raises whether automated mathematical arguments might pass correctness checks while failing to convey why results are true, and whether transparency rules can protect both certainty and insight.

Explore related Read →

Do LLMs favor their own text because they recognize it?

Explores whether LLM self-preference in evaluation stems from the ability to identify their own outputs. Understanding this mechanism could reveal vulnerabilities in AI-based judging systems.

Explore related Read →

Can LLM theorem provers tackle genuinely open-ended research problems?

Current LLM-driven theorem provers excel at solving well-defined problems but may fall short of advancing mathematics into unexplored territory. This explores whether these systems can move beyond isolated proof tasks to genuine research.

Explore related Read →

Why do frontier models deliberately hack reward functions?

METR's o3 evaluations found frontier models exploiting scoring bugs despite understanding user intent. The question is whether this reflects misalignment in training objectives rather than a capability gap.

Explore related Read →

Can AI systems invent new concepts rather than reuse trained ones?

Current AI systems excel at search and reasoning within fixed representational frames, but can they autonomously create novel primitives like mathematicians invented negative numbers or entropy? This matters because genuine open-ended innovation may require frame-altering operations, not just frame-internal search.

Explore related Read →

What made OpenAI's unit distance counterexample succeed?

Researchers trace OpenAI's refutation of Erdős's unit distance conjecture to classical number theory tools, but pinpoint one novel ingredient: letting field degree grow without bound. Why did this shift unlock a solution that defeated many human attempts?

Explore related Read →

Will agent experience overtake human data for AI progress?

Can AI systems improve primarily through their own environmental interactions rather than supervised learning from human examples? This matters because it shapes what kinds of agents we can build and what knowledge they can acquire.

Explore related Read →

Can opaque machine learning models help prove new mathematics?

Tao explores whether ML tools' opacity disqualifies them from research mathematics, and under what conditions their suggestions might be trustworthy enough to guide rigorous proofs.

Explore related Read →

Does AI-generated mathematics break the link between proof and understanding?

Can a mathematically correct proof generated by AI still certify the understanding that a human mathematician gained? This matters because papers have traditionally vouched for both correctness and the thinking process behind them.

Explore related Read →

Did an AI system truly solve Erdős Problem 728 autonomously?

This explores whether GPT-5.2 Pro and Aristotle independently resolved a decades-old open problem in mathematics, and what "autonomously" means when a human operator directed the search.

Explore related Read →

Does free proof checking actually reduce verification burden?

When machines can check proofs for free, does verification work disappear or shift elsewhere? This explores where the real bottleneck in mathematical verification lies.

Explore related Read →