Can artificial intelligence ever truly understand science?
Researchers ask whether AI can move beyond predicting outcomes to genuinely grasping the theories behind them. The question hinges on what scientific understanding actually means.
The authors argue that AI can take part in "android-assisted scientific understanding" in three ways: as a "computational microscope" that reveals properties of a system that are "otherwise difficult or even impossible to probe", as a muse supplying "new concepts and ideas that are subsequently understood and generalized by human scientists", and as an agent that gains understanding itself. In the first two, humans do the understanding. For the third they write that "we have not found any evidence of computers acting as true agents of understanding in science yet." The framing starts from an oracle that "correctly predicts" every experiment and still leaves scientists unsatisfied, because they "want to comprehend how the oracle conceived these predictions."
The criterion comes from the contextual theory of Henk de Regt and Dennis Dieks: a phenomenon can be understood if scientists "can recognise qualitatively characteristic consequences of T without performing exact calculations", where T is an intelligible theory of that phenomenon. On this view understanding sits in the scientist's ability to reason from a theory, not in the accuracy of predictions. That is what sorts the dimensions. In the first two, "the android enables humans to gain new scientific understanding"; in the third, "the machine gains understanding itself." The authors also say they surveyed dozens of scientists and collected dozens of anecdotes, but the excerpt does not reproduce them.
Set against the nearest notes, this is a graded account of understanding that differs from the mechanistic one. The mechanistic-interpretability note grades an LLM's internal organization into three tiers; this source grades the machine's role relative to a human scientist, so the two answer different questions and can be read side by side. The oracle is the case the heuristics note describes: a transformer that predicts orbital trajectories accurately without a Newtonian model is the kind of oracle the authors say a scientist would not accept. The potemkin note is a contrast. It shows that a correct explanation is not enough for understanding, while this source supplies a positive test, recognizing consequences without calculation, that the potemkin pattern can be checked against. On the source's terms, the output of the automated Mechanist instrument belongs to the first dimension, since "humans then lift these insights to scientific understanding."
The excerpt establishes nothing as verified. It uses "correct" for the oracle's predictions and "understanding" for grasping how they arise, and it describes no check by a proof assistant, an evaluator or a panel of experts. The three dimensions are argued from the philosophy of science and from anecdotes that are not shown, and the agent dimension is, in the authors' own framing, a roadmap for which they have no evidence. The implication is modest. The taxonomy is best read as a research agenda and as a way to ask, of any claimed AI result, which dimension it belongs to and who does the understanding, not as a finding that any existing system understands.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do users confuse explanation quality with actual system accuracy? Can mechanistic interpretability methods reliably reveal what models actually know?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models understand in fundamentally different ways?
Does mechanistic evidence reveal distinct tiers of understanding in LLMs—from concept recognition to factual knowledge to principled reasoning? And do these tiers coexist rather than replace each other?
parallel graded account; grades model internals where this grades the machine's role for a scientist
-
Can LLMs understand concepts they cannot apply?
Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.
contrast: shows a correct explanation is insufficient; this source supplies a positive test
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
the accurate but unexplained oracle, the case the authors say scientists reject
-
Can AI automate the discovery of how AI models work?
Whether mechanistic understanding of AI systems—traditionally manual and slow—can be accelerated through an agentic system grounded in structured knowledge and curated methods. This matters because AI development is outpacing our ability to understand it.
automated mechanism output falls under the first dimension, which needs a human to lift it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- On scientific understanding with artificial intelligence
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- ASI-Bench: At the Dawn of Artificial Superintelligence
- Open Problems in Mechanistic Interpretability
- We'll Be Arguing for Years Whether Large Language Models Can Make New Scientific Discoveries
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- Exploring the use of AI authors and reviewers at Agents4Science
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Original note title
android-assisted scientific understanding has three dimensions — microscope, muse and agent — and no true agent has been found yet