SYNTHESIS NOTE
Topics›Philosophy Subjectivity›this note

Can LLMs understand concepts they cannot apply?

Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.

Synthesis note · 2026-02-21 · sourced from Philosophy Subjectivity

The Potemkin understanding paper identifies a failure pattern that is categorically different from ordinary LLM error. When a model correctly explains an ABAB rhyme scheme, then fails to generate one, then recognizes that its generation doesn't rhyme — that triple combination is not just wrong, it is incoherent. No human with that explanation would behave that way. The combination is irreconcilable with any human cognitive pattern.

This is worth separating from other LLM failure types because the mechanism matters for diagnosis and repair:

The "Potemkin" framing (after Potemkin villages — facades with nothing behind) is precise: the model passes benchmark tests designed to detect understanding because those benchmarks test the same cognitive operations as humans. The tests only work as diagnostics if LLMs misunderstand concepts the same way humans do. But Potemkin understanding means the model can perform at the surface without the underlying integration that tests were designed to probe.

Benchmarks used to evaluate LLMs are also used to evaluate people. They are valid tests only if LLMs fail in human-compatible ways. Potemkin understanding shows that this assumption fails — LLMs can fail in ways that no human cognitive model predicts.

The three-domain evidence (literary techniques, game theory, psychological biases) shows this is not domain-specific. Across domains: near-perfect explanation accuracy, significant application failure, model recognition of failure. The incoherence is stable.

The "computational split-brain syndrome" diagnosis. "Comprehension Without Competence" provides the architectural analysis underlying Potemkin understanding. Through controlled experiments, the authors demonstrate that instruction and action pathways are geometrically and functionally dissociated — a phenomenon they term computational split-brain syndrome. The failure is not in knowledge access but in computational execution. LLMs function as powerful pattern completion engines but lack the architectural scaffolding for principled, compositional reasoning. This diagnosis also clarifies why mechanistic interpretability findings may reflect training-specific pattern coordination rather than universal computational principles. The geometric separation between instruction and execution pathways represents a structural limitation, not a knowledge limitation.

The Explain-Query-Test (EQT) framework provides direct empirical measurement of the explanation-comprehension gap. In EQT, a model (1) generates an explanation of a topic, (2) generates question-answer pairs from that explanation, and (3) answers those same questions without access to its own explanation. The finding: models consistently fail questions derived from their own explanations. The EQT gap correlates strongly with MMLU-PRO benchmark performance — making EQT a benchmark-free evaluation method that uses only the model's own outputs as ground truth. Critically, the gap is domain-specific: biology and psychology (domains where models initially perform well) show the largest EQT drops, while law and engineering (lower baseline) show smaller drops. This suggests Potemkin understanding is worst precisely where surface performance is highest — a counterintuitive result that demands explanation. High benchmark performance may mask explanation-comprehension disconnection rather than reveal genuine understanding.

Inquiring lines that read this note 216

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we detect and account for LLM involvement in academic writing? Can language models reason beyond surface pattern matching? Why do models reveal hidden associations despite concealment attempts? Why do language models hallucinate and how can we prevent it? Can LLMs distinguish between linguistic form and semantic meaning? What limits language model accuracy in evaluating ideas? How does RLHF training shape models to prioritize agreement over accuracy? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do LLM research ideation systems generate novelty but lack diversity? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Do language models reason through disagreement or only accommodate it? What explains the gap between benchmark scores and true reasoning capability? Should models ask for clarification when facing ambiguous or under-specified information? Why do language models fail at sustained therapeutic relationships despite understanding techniques? Can AI systems discover fundamental improvements to their own architectures? What prevents language models from performing systematic logical reasoning? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can models develop genuine introspective capability, or only mimic it? Why do language models struggle to implement user intent accurately from prompts? How does model capacity affect learning performance on diverse downstream tasks? Can mechanistic interpretability methods reliably reveal what models actually know? How does fine-tuning trade off accuracy against reasoning quality? Is embodied interaction necessary for language meaning and agency? How do training data quality and composition affect downstream model performance? How can we reduce inherent biases in LLM-based evaluation judges? How reliably can language models perform causal versus temporal reasoning? Can reasoning models use reflection to correct their initial outputs? Why do training associations persist despite contradictory contextual information? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What distinguishes genuine communicative competence from surface language performance? Why don't better reasoning capabilities improve theory of mind performance? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Why do standard evaluation practices obscure safety-critical AI failures? What causes coordination failures in multi-agent language model systems? How can persistent memory architectures preserve information across ultra-long contexts? How do users confuse explanation quality with actual system accuracy? When do simpler collaborative filtering approaches outperform complex LLM recommenders? How do neural networks learn compositional structure from training? Can smaller specialized models match frontier models on key metrics? What makes reasoning traces effective supervision even when they're incorrect? Can we trust AI-generated mathematical proofs without understanding them? How should retrieval strategies adapt to multi-step reasoning demands? How do hallucinated citations emerge in AI scholarly output?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
21 direct connections · 220 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

potemkin understanding is a distinct failure mode where correct explanation combined with failed application is incoherent not merely wrong