SYNTHESIS NOTE
Topics›MechInterp›this note

Do language models understand in fundamentally different ways?

Does mechanistic evidence reveal distinct tiers of understanding in LLMs—from concept recognition to factual knowledge to principled reasoning? And do these tiers coexist rather than replace each other?

Synthesis note · 2026-04-18 · sourced from MechInterp

This paper synthesizes mechanistic interpretability findings into a philosophical framework that moves beyond the binary "does AI understand?" debate. The framework proposes three hierarchical tiers:

Tier 1: Conceptual understanding — arises when a model forms "features" as directions in latent space that unify diverse manifestations of a single entity or property. This is the representational foundation: the model has learned that different surface forms connect to the same underlying concept. MI evidence: SAE features, linear probing, representation geometry studies all demonstrate this.

Tier 2: State-of-the-world understanding — arises when the model learns contingent factual connections between features and dynamically tracks changes. "Michael Jordan is a basketball player" is not just a high-probability string but a reflection of an internal model linking the Michael Jordan concept to the basketball player concept. This goes beyond association to structured knowledge representation.

Tier 3: Principled understanding — arises when the model discovers compact "circuits" that connect facts via general rules rather than memorizing each fact individually. This is the shift from knowing that to knowing why. The grokking literature provides the clearest evidence: models that transition from memorization to generalization develop circuits implementing actual algorithmic rules (e.g., modular addition via Fourier transforms).

The critical insight is that higher-tier mechanisms coexist with lower-tier heuristics rather than replacing them. A model can have principled understanding of arithmetic in one circuit while relying on pattern-matching heuristics in another. This heterogeneity means understanding is not a single binary property but a patchwork: principled in some domains, merely conceptual in others, and purely heuristic in yet others.

This has direct implications for trust and deployment. The fact that a model demonstrates principled understanding in one domain gives no guarantee that it operates at the same tier in adjacent domains. The coexistence of understanding tiers also explains why models can be simultaneously impressive and brittle: the principled circuits work reliably, but the heuristic patches fail unpredictably.

Inquiring lines that read this note 84

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can mechanistic interpretability methods reliably reveal what models actually know? What prevents LLMs from applying their reasoning knowledge to improve outputs? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Can AI systems discover fundamental improvements to their own architectures? Can language models reason beyond surface pattern matching? What prevents language models from performing systematic logical reasoning? Can models develop genuine introspective capability, or only mimic it? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Should models ask for clarification when facing ambiguous or under-specified information? Do language models reason through disagreement or only accommodate it? Is embodied interaction necessary for language meaning and agency? What distinguishes genuine communicative competence from surface language performance? Can LLMs distinguish between linguistic form and semantic meaning? What limits language model accuracy in evaluating ideas? What explains the gap between benchmark scores and true reasoning capability? How should retrieval strategies adapt to multi-step reasoning demands? How do neural networks learn compositional structure from training? Can base models hide emergent misalignment through alignment training? Can minimal training unlock latent reasoning already present in base models? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Why don't better reasoning capabilities improve theory of mind performance? How do curriculum design and feedback approaches affect model learning? How do users confuse explanation quality with actual system accuracy? How reliably can language models perform causal versus temporal reasoning?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 168 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

mechanistic interpretability evidence supports three hierarchical varieties of LLM understanding — conceptual then state-of-world then principled — each tied to a distinct computational organization