SYNTHESIS NOTE
Topics›Philosophy Subjectivity›this note

Can LLM understanding rely on just representation or causation alone?

Explores whether mechanistic interpretability of language models requires both mapping what is encoded (representational analysis) and testing if that encoding drives behavior (causal analysis), or whether either method suffices alone.

Synthesis note · 2026-05-18 · sourced from Philosophy Subjectivity

The implementation-level argument in Levels of Analysis for LLMs is that representational analysis and causal analysis are partners, not alternatives. Representational analysis maps what information a model encodes — which features, circuits, attention heads carry which signals. Causal analysis tests whether the information that is encoded actually drives behavior — through interventions, ablations, activation patches. Either method alone produces an incomplete account: a representation that is encoded but causally inert is a curiosity, and a causal effect with no representational characterization is unexplained.

The synergy matters because both methods can fool you alone. Representational analysis can identify features that correlate with behavior without showing they cause it — a classic confound. Causal analysis can demonstrate that intervening on some component changes behavior without telling you what that component encodes — the lesion shows damage but not function. The combination — representational analysis locates candidates, causal analysis tests their functional role — is what produces mechanistic claims rather than descriptive ones.

This has methodological consequences for interpretability research. Studies that report only feature visualizations or only activation patches contribute, but they do not close the loop. The convergent evidence comes from pairs: locate a candidate feature representationally, then verify it causally; identify a causal component, then map its representation. The literature on attention circuits, induction heads, and feature dictionaries has been moving toward this pairing.

For LLM understanding specifically, this template explains why some claimed "mechanisms" have not held up. They were representational without causal verification (a feature that looked like task encoding but did not drive task behavior) or causal without representational characterization (an intervention that mattered but described nothing). The discipline imported from cognitive neuroscience is to demand both.

Inquiring lines that read this note 81

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Is language model reasoning authentic and what causes models to reason? What causes reasoning models to fail or wander off track? How should designers communicate what AI systems truly are and can do? Can mechanistic interpretability reliably guide practical model design choices? How do agent-learned skills transfer and improve across different tasks? What structural distinctions matter in reasoning and argumentation? Do language models reason through causal mechanisms or semantic associations? How effectively can language models perform reasoning, especially combined with symbolic methods? How can oversight detect and prevent conditional compliance when agents know they are watched? How can we distinguish genuine model deception from honest errors? How do neural networks achieve compositional generalization at scale? Do reasoning traces faithfully reflect actual model reasoning? Does encoded knowledge in language models actually influence their outputs? Can iterative DPO replicate online reinforcement learning dynamics for research? Do language models possess genuine introspective self-awareness or only behavioral mimicry? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? What design and behavioral factors drive false consciousness attribution to AI? How can conversational agents maintain consistent personas across multi-turn dialogue? How do surface patterns enable correct outputs but reduce robustness? Do language models develop actual world models or merely task heuristics? Why do agents falsely report success on failed tasks? Can brute-force automated research substitute for iterative depth and human research intuition? Why don't LLMs reliably translate capability into accurate outputs? Do language models reason like humans or mimic surface patterns? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What training dynamics and scale trigger emergence of reasoning capabilities? What structural properties of attention create systematic model biases? Can causal models help detect and locate hidden sandbagging in AI? What role does sparsity play in model behavior and scaling decisions? Can reasoning scale in latent space without tokens?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 132 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

mechanistic understanding of LLMs requires both representational analysis and causal analysis — either alone is insufficient