INQUIRING LINE

An AI can predict results without understanding them, so what does a scientist need before they feel they've really explained something?

What makes a scientist satisfied with an explanation versus just a prediction?

This explores what separates an explanation from an accurate prediction in science, and why scientists tend to want the first even when they already have the second. The corpus has no study of scientists' own sense of satisfaction, so it answers from the side: through work on AI models that predict well without understanding anything.


This explores the gap between knowing *that* something will happen and understanding *why*, and what scientists treat as the test for the second. The corpus has no study of how scientists themselves feel about explanations. It does contain a sharp set of cases where AI delivers predictions with no understanding behind them, and those cases show what explanation is supposed to add.

Start with how good prediction alone has become. Fine-tuned language models beat neuroscience experts at guessing which experimental results actually happened. The same pattern-blending habit that produces hallucinations when you ask about the past turns into real foresight when you ask about the future Can LLMs predict novel scientific results better than experts?. Yet a framework for AI in science finds that AI does well as a tool and as a source of ideas, but has not yet shown understanding as an independent agent Can artificial intelligence ever truly understand science?. So beating the experts at prediction does not get a model counted as understanding the science. Something else is being measured.

One strong candidate is that an explanation has to hold up under "what if things were different?" Work on decision-making shows that a model can be accurate on average yet systematically wrong in exactly the situations where the decision matters Why do accurate predictions lead to poor decisions?. Explanations of AI behavior show the same weakness from the other side. Explanations that people judge correct and coherent fail to predict what the model does once you change the inputs, and RLHF makes them more convincing without making them more accurate Can LLM explanations actually help humans predict model behavior?. An explanation that only feels satisfying is a trap. A good one should let you predict cases you haven't seen yet. Explanations also help people catch AI errors only when they argue both for and against the answer Do explanations actually help users spot AI mistakes?. That echoes how scientists actually trust a theory: by knowing what would prove it wrong.

The most useful idea here comes from philosophy of science. Opacity, not being able to see why a model gives its output, matters chiefly at the point of *justification*, not *discovery* Can opaque models guide discovery without needing interpretation?. An opaque model's prediction can legitimately point scientists toward something new. Scientists are satisfied only once they've built a theory that passes their field's own standards, without relying on the model. In other words, prediction can open the door, but explanation is what a field agrees to sign off on.

That signing-off is social, which is the part most people don't expect. The meaning of an explanation is set by groups observing and interpreting one another, not inside a single person's head Where does the meaning of an AI explanation actually come from?. Whether an explanation works also depends on who presents it, how it is framed, and who receives it What if XAI is fundamentally a communication problem?. Scientific discovery itself becomes much easier to forecast once you model *the scientists*, meaning who collaborates with whom and who has expertise in which materials, rather than just the content of papers Can predicting scientists improve discovery forecasts?. So the satisfaction you're asking about may be less a private 'aha' than a community's verdict that a story can be taught, defended and extended.


Sources 9 notes

Can LLMs predict novel scientific results better than experts?

BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.

Can artificial intelligence ever truly understand science?

A framework distinguishes three roles for AI in science: as a computational tool, a source of ideas, and as an independent agent. Evidence shows AI succeeds in the first two but has not yet achieved understanding in the third role.

Why do accurate predictions lead to poor decisions?

Research formalizes necessary and sufficient conditions for predictive models to support optimal decisions. A model can predict accurately on average yet systematically mispredict in decision-critical states.

Can LLM explanations actually help humans predict model behavior?

Explanations that humans judge as correct and coherent fail to predict model behavior on counterfactuals. RLHF optimization improves how convincing explanations seem without improving their actual predictive accuracy, leaving users confident but wrong.

Do explanations actually help users spot AI mistakes?

Reasoning traces and post-hoc explanations increase user acceptance of AI answers regardless of correctness, engendering false trust. Only dual explanations presenting arguments for and against the answer genuinely help users distinguish correct from incorrect outputs.

Show all 9 sources
Can opaque models guide discovery without needing interpretation?

Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.

Where does the meaning of an AI explanation actually come from?

Drawing on Luhmann's multi-layer cybernetics, AI explanation meaning is constituted at the social-group level through layered observations of observations, not produced inside dyadic human-AI dialogue. Lab-tested explanations stripped of social context will not predict real-world effectiveness.

What if XAI is fundamentally a communication problem?

Explanation quality is not intrinsic to the explanation itself but depends on the rhetorical situation: who presents it, how it is framed, and what role the recipient plays. Evaluations that ignore this triad measure only a narrow slice of real-world effectiveness.

Can predicting scientists improve discovery forecasts?

Random walks over hypergraphs of papers, materials, and authors forecast discoveries 43% more precisely than content-only models, especially when literature is sparse. The mechanism simulates plausible scientific inference steps like collaboration and material expertise.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.