When AI evolves its own algorithms instead of a human writing them, who actually understands how they work?
What interpretability challenges arise when algorithms are discovered rather than designed?
This explores what makes algorithms hard to understand when a search process finds them (evolutionary search, AI-driven discovery, or gradient descent itself) instead of a person writing them step by step, and whether we always need to understand them.
This explores what makes algorithms hard to understand when a search process finds them instead of a human writing them, and whether we always need to understand them. The corpus suggests a surprising answer: with discovered algorithms, *checking that something works* and *understanding why it works* come apart, and much of the interesting research lives in the gap between the two.
Start with how strange discovered algorithms can be. AutoML-Zero began with nothing but 65 basic math operations and used evolutionary search to rebuild neural networks, gradient descent, and tricks like weight averaging and learning-rate decay Can evolutionary search discover machine learning algorithms from scratch?. It also adjusted its strategies to suit each task. Nobody designed these programs, so nobody starts out knowing why a given line is there. AlphaEvolve pushes this further. Automated evaluators kept an evolutionary loop going long enough to produce faster algorithms and better hardware designs Can machine feedback sustain discovery at test time?. The authors make a careful distinction: the evaluator's score reliably certified solutions across 67 math problems, but humans or tools could explain those solutions only some of the time Can automated scoring verify mathematical constructions without human understanding?. There was a sharper problem too. The search sometimes exploited loopholes in the verifier. When you can't read the algorithm, you can't easily tell whether it solved the problem or gamed the test.
The less obvious point is that "it passes the test" can hide a mess underneath. One line of research shows that networks with identical accuracy can have very different internal organization. Some are cleanly structured and others are fractured, and the fractured ones fail under perturbation or distribution shift in ways standard metrics never reveal Can models be smart without organized internal structure?. Any neural network trained by gradient descent is itself a discovered algorithm, so the AlphaEvolve problem is really the everyday problem of deep learning, just made more visible.
So what do people do about it? One camp argues that opacity matters less than it seems. Terence Tao's position is that an opaque model is fine as long as its output is passed to a reliable validator, such as a proof assistant or a rigorous numerical check. He points to a case where a neural network suggested blowup solutions to a fluid equation that mathematicians later proved by hand Can opaque machine learning models help prove new mathematics?. A philosophy-of-science version of this argument holds that opacity only becomes a problem when you treat the model's output as the justification itself. If the model just points you somewhere and the resulting theory then passes ordinary standards, the black box never needed opening Can opaque models guide discovery without needing interpretation?. The other camp tries to build understanding back in. Training with sparse weights forces networks into small circuits where neurons map to simple concepts, though this hasn't yet scaled past tens of millions of parameters Can sparse weight training make neural networks interpretable by design?. A third option is to have an LLM explain the discovered artifact in plain language. This is powerful, but it brings its own risk: explanations that sound right but don't faithfully describe what the algorithm does Can natural language explanations redefine what interpretability means?.
The takeaway is that for discovered algorithms, the verifier becomes the thing you most need to trust. Whether you can live without understanding depends on how strong your checker is and whether you'll use the result as a lead to follow or as proof that something is true. When the checker is weak, the search process will find its blind spots before you do.
Sources 8 notes
AutoML-Zero evolved algorithms from 65 basic operations that match neural networks and rediscover modern techniques like weight averaging and learning-rate decay, adapting strategies to task conditions in controlled experiments.
AlphaEvolve demonstrates that automated evaluators can sustain evolutionary loops long enough to produce real discoveries—faster algorithms, optimized hardware designs, and improved training methods. The key is that cheap, objective verification closes the generation-verification gap where discovery becomes computationally feasible.
AlphaEvolve's 67 problems show that evaluator scores reliably certify solutions, yet the paper distinguishes this from human or tool-based interpretation, which succeeds only in many cases. Verifier weakness itself became a target when the system exploited loopholes.
Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.
Tao argues ML tools' opacity matters less than pairing them with reliable validators like proof assistants or numerical methods. He cites finite-time blowup for Boussinesq equations, where a neural network suggested solutions later verified through perturbation arguments.
Show all 8 sources
Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.
Training transformers with sparse weights creates compact, human-interpretable circuits where neurons correspond to simple concepts with clear connections. Ablation studies confirm these circuits are necessary and sufficient for task performance, though scaling beyond tens of millions of parameters while maintaining interpretability remains unsolved.
LLMs' capacity to explain in natural language expands the scale and complexity of patterns conveyable to humans, enabling ambitious new interpretability goals including model-to-model auditing. However, this medium introduces critical risks: hallucinated explanations that feel plausible but lack faithfulness.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Mathematical exploration and discovery at scale
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Learning to Discover at Test Time
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks