INQUIRING LINE

A model can predict well on average yet be badly wrong in the exact moments where your decision actually hinges on it.

When does prediction without causal mechanism fail to support good decisions?

This explores when a model that predicts well, but has no account of *why* things happen, still leads people to bad choices, and what the corpus says about where that gap opens up.


This explores when a model that predicts well, but has no account of *why* things happen, still leads people to bad choices. The short answer from the corpus is that prediction tends to break down at the moment you act on it. A model can be accurate on average and still be systematically wrong in the few situations where the decision actually matters. Why do accurate predictions lead to poor decisions? spells out the conditions under which a predictive model supports good decisions. Being good at fitting the data is not one of them. When a model is tuned for overall accuracy, its errors can pile up in the cases that decide which option you pick, and a high average score hides that.

A second failure point is spurious correlation, where a pattern holds in the training data for reasons that have nothing to do with cause. If a model has only learned that two things tend to occur together, its prediction tells you little about what happens when you step in and change one of them. One response in the corpus is to stop asking the language model to do the causal reasoning at all. Can separating causal models from language models improve reasoning? keeps an explicit causal model that can be revised as evidence comes in, and uses the LLM only to translate that model's outputs into language. Interpretability research makes the same point about studying models themselves. Finding internal features that go along with a behavior does not show they cause it. You have to intervene and watch what changes, as Can LLM understanding rely on just representation or causation alone? argues.

The less obvious point is that prediction without mechanism doesn't always fail. It depends on what you use the prediction for. Can opaque models guide discovery without needing interpretation? separates using a black-box prediction to *point you somewhere* (which experiment to run, which hypothesis to test) from using it as the *justification* for a conclusion. In the scientific cases it describes, opaque predictions led to theories that then held up under the field's own standards of evidence. The trouble starts when the prediction itself is treated as the reason for the decision, with no independent check afterward.

When you can't get the mechanism, one partial safeguard is a model that knows when it doesn't know. Can models learn to abstain when uncertain about predictions? found that small models trained to report honest confidence, and to abstain when unsure, matched models ten times their size at forecasting how conversations would unfold. That doesn't supply causal understanding, but it can flag the cases where acting on the prediction would be risky. The corpus also has a warning about overcorrecting. Removing misleading cues can make performance *worse* when the real difficulty is weighing signals that conflict, as Why does removing spurious cues sometimes hurt model performance? shows. Causal models aren't a complete replacement either, since they leave out the associative, analogical and emotional parts of how people actually reason (Can causal models alone capture how humans actually reason?).

The takeaway: ask what a prediction is about to be used for, not only how accurate it is. If it's a lead to investigate, mechanism can come later. If it's the reason you're acting, especially in rare situations where an average-case score tells you little, you need either a causal account or a model honest enough to say it doesn't know.


Sources 7 notes

Why do accurate predictions lead to poor decisions?

Research formalizes necessary and sufficient conditions for predictive models to support optimal decisions. A model can predict accurately on average yet systematically mispredict in decision-critical states.

Can separating causal models from language models improve reasoning?

Causal Reflection separates causal reasoning into a formal dynamic model with a Reflect mechanism for revision, relegating the LLM to structured inference and language rendering. This architecture sidesteps asking LLMs to perform causal reasoning directly, addressing both spurious-correlation failures and RL's explanation gap.

Can LLM understanding rely on just representation or causation alone?

Research shows that representational analysis alone identifies correlates without proving causation, while causal analysis alone demonstrates effects without explaining function. Only paired methodology—locating candidates representationally then verifying causally—produces genuine mechanistic understanding rather than descriptive claims.

Can opaque models guide discovery without needing interpretation?

Deep learning models can guide discovery through opaque outputs without interpretation because justification applies to the resulting theory, not the model. Two cases show accurate predictions leading to theories that pass disciplinary standards independent of model understanding.

Can models learn to abstain when uncertain about predictions?

Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.

Show all 7 sources
Why does removing spurious cues sometimes hurt model performance?

Removing spurious cues degrades performance in heuristic override tasks, opposite to shortcut learning predictions. The failure mode is integrating conflicting signals rather than ignoring distractors—a frame problem, not feature selection.

Can causal models alone capture how humans actually reason?

Causal belief networks excel at modeling causal reasoning but cannot represent associative links, analogical mappings, or emotion-driven belief shifts. The GenMinds framework itself acknowledges this as a tractable starting point rather than a complete theory.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.