INQUIRING LINE

FunSearch says its programs tend to be readable, but its scorer only checks whether they work, not whether anyone understands them.

Why does FunSearch claim interpretability without measuring human comprehension?

This explores why FunSearch, a system that pairs a language model with an automatic scorer to discover new programs, can call its results interpretable when nobody checked whether people actually understood them, and what the rest of the collection says about the gap between readable output and real understanding.


This explores why FunSearch calls its results 'interpretable' when nobody tested whether people could understand them. The short answer is that FunSearch never tries to measure comprehension. Its proof that a program works comes entirely from an automatic scoring function that runs on every candidate. The interpretability claim is a softer add-on: the programs are only said to *tend* to be readable, and there is no evidence that anyone studied whether humans grasped them How does FunSearch actually verify its discovered programs?. So the two claims rest on different footing. Correctness is measured. Interpretability is assumed because the output is code, and code is something a person could in principle read.

That split follows from how the system is built. Language models are good at proposing candidates and bad at judging how good those candidates really are, so discovery systems pair them with an outside judge Can language models reliably judge their own candidate quality?. FunSearch's scorer is that judge. It can tell whether a program performs well. It cannot tell whether a human understands why. Nothing in the loop rewards clarity, so any readability comes as a side effect.

The broader interpretability research in the collection shows why 'it's readable' and 'it's understood' are different claims. One line of work replaces attention heads with synthesized programs that match their behavior 99% of the time, and calls this 'formal, testable interpretability' Can executable programs replace transformer attention heads?. Even there, 'testable' means you can check that the program behaves the same way, not that a person learns something from reading it. A related note argues that real mechanistic understanding needs two things: finding a candidate explanation, and then proving it is what actually causes the behavior. Either step alone gives you a description, not understanding Can LLM understanding rely on just representation or causation alone?. FunSearch has the second half, a verified effect. It does not have the first half, a demonstrated explanation of why the program works.

The surprising point is that a readable surface can actively mislead. Natural-language explanations make it possible to explain far more to people, but they can also produce explanations that sound right without matching what the model actually does Can natural language explanations redefine what interpretability means?. Models themselves show this split: they can explain a concept correctly and still fail to apply it Can LLMs understand concepts they cannot apply? Can language models understand without actually executing correctly?. If being able to explain something doesn't prove a model understands it, then being able to read code doesn't prove a human does either. RecExplainer offers a model for treating both as separate goals. It aims for explanations that are faithful to the system and also understandable to users Can LLMs explain recommenders by mimicking their internal states?.

What would measuring comprehension look like? The collection hints at an approach. Chatbot Arena shows that large numbers of human judgments can be collected cheaply and line up with expert opinion Can crowdsourced votes reliably rank language models?. A similar study could ask people to predict what a FunSearch program does or to modify it. The collection has only one note directly about FunSearch, so this is an inference from nearby work and does not report what its authors intended. The pattern itself is common and worth spotting elsewhere: when a paper measures one property and assumes another, look at which claim got an experiment.


Sources 9 notes

How does FunSearch actually verify its discovered programs?

FunSearch's verification relies on a scoring function applied to each candidate program, while interpretability claims are only hedged as a tendency. The excerpt shows no evidence that humans actually understood the discovered programs.

Can language models reliably judge their own candidate quality?

LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.

Can executable programs replace transformer attention heads?

Program synthesis recovers executable code matching 99% of attention head behavior. Substituting the best-fit 30–40% of heads with their synthesized programs preserves QA ability, offering formal, testable interpretability instead of natural-language summaries.

Can LLM understanding rely on just representation or causation alone?

Research shows that representational analysis alone identifies correlates without proving causation, while causal analysis alone demonstrates effects without explaining function. Only paired methodology—locating candidates representationally then verifying causally—produces genuine mechanistic understanding rather than descriptive claims.

Can natural language explanations redefine what interpretability means?

LLMs' capacity to explain in natural language expands the scale and complexity of patterns conveyable to humans, enabling ambitious new interpretability goals including model-to-model auditing. However, this medium introduces critical risks: hallucinated explanations that feel plausible but lack faithfulness.

Show all 9 sources
Can LLMs understand concepts they cannot apply?

Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.

Can language models understand without actually executing correctly?

Large language models can articulate correct principles but systematically fail to apply them due to dissociated instruction and execution pathways. The 87% accuracy in explanations versus 64% in actions reveals this is not knowledge deficit but structural disconnect.

Can LLMs explain recommenders by mimicking their internal states?

RecExplainer trains LLMs via three alignment methods: behavior (mimicking outputs), intention (incorporating neural embeddings), and hybrid (combining both). The hybrid approach produces explanations that are simultaneously faithful to the target model and intelligible to users by balancing internal-state inspection with human-readable reasoning.

Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.