Can LLMs understand concepts they cannot apply?
Explores whether large language models can correctly explain ideas while simultaneously failing to use them—and whether that combination reveals something fundamentally different from ordinary mistakes.
The Potemkin understanding paper identifies a failure pattern that is categorically different from ordinary LLM error. When a model correctly explains an ABAB rhyme scheme, then fails to generate one, then recognizes that its generation doesn't rhyme — that triple combination is not just wrong, it is incoherent. No human with that explanation would behave that way. The combination is irreconcilable with any human cognitive pattern.
This is worth separating from other LLM failure types because the mechanism matters for diagnosis and repair:
- Ordinary errors (fabrication, factual mistakes) — the model lacks information or generates plausible-but-false continuations. Fix: better retrieval, grounding, training data.
- Surface generalizations — the model learned correlations that worked in training but don't generalize structurally. Fix: better training curriculum, structural probing.
- Potemkin understanding — the model can produce the explanation and fails to apply it and recognizes the failure. This combination implies that explanation-generation and concept-application are functionally disconnected. No single epistemic fix addresses both.
The "Potemkin" framing (after Potemkin villages — facades with nothing behind) is precise: the model passes benchmark tests designed to detect understanding because those benchmarks test the same cognitive operations as humans. The tests only work as diagnostics if LLMs misunderstand concepts the same way humans do. But Potemkin understanding means the model can perform at the surface without the underlying integration that tests were designed to probe.
Benchmarks used to evaluate LLMs are also used to evaluate people. They are valid tests only if LLMs fail in human-compatible ways. Potemkin understanding shows that this assumption fails — LLMs can fail in ways that no human cognitive model predicts.
The three-domain evidence (literary techniques, game theory, psychological biases) shows this is not domain-specific. Across domains: near-perfect explanation accuracy, significant application failure, model recognition of failure. The incoherence is stable.
The "computational split-brain syndrome" diagnosis. "Comprehension Without Competence" provides the architectural analysis underlying Potemkin understanding. Through controlled experiments, the authors demonstrate that instruction and action pathways are geometrically and functionally dissociated — a phenomenon they term computational split-brain syndrome. The failure is not in knowledge access but in computational execution. LLMs function as powerful pattern completion engines but lack the architectural scaffolding for principled, compositional reasoning. This diagnosis also clarifies why mechanistic interpretability findings may reflect training-specific pattern coordination rather than universal computational principles. The geometric separation between instruction and execution pathways represents a structural limitation, not a knowledge limitation.
The Explain-Query-Test (EQT) framework provides direct empirical measurement of the explanation-comprehension gap. In EQT, a model (1) generates an explanation of a topic, (2) generates question-answer pairs from that explanation, and (3) answers those same questions without access to its own explanation. The finding: models consistently fail questions derived from their own explanations. The EQT gap correlates strongly with MMLU-PRO benchmark performance — making EQT a benchmark-free evaluation method that uses only the model's own outputs as ground truth. Critically, the gap is domain-specific: biology and psychology (domains where models initially perform well) show the largest EQT drops, while law and engineering (lower baseline) show smaller drops. This suggests Potemkin understanding is worst precisely where surface performance is highest — a counterintuitive result that demands explanation. High benchmark performance may mask explanation-comprehension disconnection rather than reveal genuine understanding.
Inquiring lines that read this note 216
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we detect and account for LLM involvement in academic writing?- Why do some LLM clusters cite broader psychology than others?
- Can knowledge density explain why LLM writing feels coherent but fatiguing?
- What structural barriers prevent LLMs from making evaluative judgments about writing?
- Can LLMs infer situational context the way humans do pragmatically?
- How do fixed pragmatic templates prevent models from understanding context?
- Why do language models fail at planning despite understanding strategies?
- Why do language models fail when semantic content is stripped away?
- Why do LLMs produce semantically acceptable but pragmatically disengaged responses?
- Can explicit connectives compensate for missing intentional tracking in LLMs?
- Can LLMs improve at metaphor if they handle decoupled semantics better?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why do large language models still have systematic blind spots with complex structures?
- Can LLMs infer implicit meaning without surface linguistic markers?
- Why do LLMs fail at implicit elements in literary and poetic text?
- Can complexity-stratified testing reveal whether LLMs understand grammatical structure?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Can LLMs translate between natural language and formal logic faithfully?
- Do metaphors work by decoupling meaning from linguistic associations?
- Can LLMs identify implicit metaphoric mappings that require pragmatic inference?
- Can LLM semantic representations exist without causally influencing their generation output?
- How do LLMs handle false presuppositions embedded in user questions?
- Why does LLM compression eliminate causal grounding in conceptual representations?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Can LLMs compute how presuppositions project through embedded clauses?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Can we use LLM language without adopting LLM assumptions?
- How do LLMs lose information when translating natural language to formal logic?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
- How faithful are natural language explanations from LLMs really?
- Can language models translate theorems faithfully without semantic loss?
- Can LLMs decode their own hidden activations into natural language?
- Why do LLMs miss new scientific ideas before they enter formal literature?
- What happens when LLMs analyze literary irony that relies on understatement?
- What cognitive capacities do LLMs actually lack that commentary assumes they have?
- Can LLMs distinguish ethical cases that differ only in critical nouns?
- How does the inability to manage ambiguity undermine literary analysis tasks?
- How much does question framing affect LLM accuracy on knowledge tasks?
- How much semantic meaning survives when LLMs paraphrase poetry and literary text?
- Can LLMs recognize rhetorical devices they cannot actually produce themselves?
- Can LLMs distinguish stylistic patterns that carry meaning from mere convention?
- Is the distinction between pretense and realization meaningful for LLMs?
- Why do LLM stories over-explain themes and favor single-track plots?
- Why does LLM fluency create false perceptions of professional standing and expertise?
- Does LLM vocabulary become the cultural lexicon for how we think?
- Do LLMs lack evaluative capacity or only default taste and stance?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Why do NLP benchmarks systematically exclude ambiguous test cases from evaluation?
- Why do NLP benchmarks exclude ambiguous instances from evaluation?
- Is interpretive multiplicity a bug in language or a feature?
- How do rare linguistic registers differ from conceptually complex examples?
- Why do large language models fail at temporal reasoning in complex legal cases?
- Why do LLMs excel at generation but struggle with evaluation?
- Does more thinking always help large language models or sometimes hurt?
- Do LLMs struggle more with semantic accuracy than syntactic correctness across domains?
- What distinguishes entity errors from relation errors in LLM output?
- Can language models accurately evaluate the quality of their own ideas?
- Why do NLP benchmarks hide LLM failures in ambiguity handling?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Why do standard NLP benchmarks hide the most critical language limitations?
- Why do language models fail at understanding ambiguous or complex requirements?
- How does the pretraining distribution shape what LLMs find hard?
- Can LLMs reliably audit other language models for errors?
- How do model compression biases differ from human conceptual representation strategies?
- Does LLM miscalibration cause failures in clinical information extraction?
- Why do LLMs excel at isolated tasks but fail at integrating components across long texts?
- Why do untrained LLMs default to rigid and sycophantic editing rules?
- How does LLM hallucination risk manifest in knowledge graph construction?
- Where do LLMs succeed at generation but struggle with evaluation?
- Why do LLM personas struggle with specificity in specialized domains like law?
- What specific execution barriers do LLM ideas encounter most frequently?
- Can output-layer corrections fix fundamental cultural representation deficits in LLMs?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- Why does LLM knowledge fail to influence their actual outputs?
- Can LLMs explain concepts correctly while failing to use them?
- What causes LLMs to ignore unstated constraints they know about?
- How do LLMs compress specific expert knowledge into median abstraction?
- Can pruning half of LLM layers affect knowledge retrieval performance?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- Where do LLMs fail as knowledge systems compared to humans?
- What internal mechanisms explain LLM reasoning and representation limits?
- Why can LLMs identify argument structure but not check warrants?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- Why can't LLMs reason from first principles or initial commitments?
- Why do LLMs explain evidence accurately while missing its implications?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Why can LLMs interpret formal logic better than they generate it?
- Do LLMs fail exploration because of context integration or computational limitations?
- What data presentation structures enable LLMs to learn decision-making from examples?
- Can training procedures fix LLM accommodation of false presuppositions?
- Can LLMs improve at simple deduction through different training approaches?
- How can a model explain something correctly yet fail to apply it?
- How does an instruction-following LLM activate latent retrieval knowledge?
- How does the LLM Fallacy prevent users from noticing cognitive debt accumulating?
- Why do experts experiencing the LLM Fallacy fail to develop custodian skills?
- Why do backward-looking benchmarks underestimate LLM scientific value?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- What concrete problems do LLMs solve at the computational level?
- What implicit knowledge about catalogs do LLMs learn from ranking signals alone?
- What latent mechanisms do LLMs use when they cannot execute iterative methods?
- What role do model-based critics play in validating LLM plans?
- Why do LLMs choose incorrect edits despite understanding the task?
- How do knowing and doing diverge in LLM decision-making?
- Can surface-level correctness hide failures in structural learning by LLMs?
- Why do LLMs struggle more when only numerical values change?
- Can irrelevant information reliably expose the limits of LLM reasoning?
- What structural framework prevents LLM explanations from becoming just plausible fiction?
- How do LLM explanations diverge from actual internal reasoning?
- Why do LLMs reason fluently about causality but lack causal rigor?
- What capability boundary exists in LLM prediction of effect sizes?
- Do LLMs detect harmful concepts before they influence model outputs?
- Why don't LLM explanations predict what models would actually do?
- What barriers prevent experts from specifying concepts for LLM extraction?
- How should LLM abstraction tools be evaluated without manual labeling?
- Why do people misinterpret or misuse LLM outputs in practice?
- What levels of understanding about LLM knowledge representation can automated systems reliably extract?
- Can LLMs give correct answers without making those answers understandable to users?
- Why does chain-of-thought reasoning alone not fix LLM performance with users?
- Why do LLMs generate ideas that sound novel but fail during execution?
- Why does LLM research ideation collapse into low diversity despite high novelty?
- What makes a novel research idea practically infeasible for implementation?
- Why do LLMs generate novel ideas but lack evaluative commitment?
- Do LLMs generate more novel ideas than they can evaluate?
- Why do LLMs generate novel ideas but struggle to evaluate them?
- Can LLMs generate more novel research ideas than human experts?
- What makes a problem instance unfamiliar to a language model?
- Why do text-to-image models fail at composing multiple concepts together?
- What happens when formal languages satisfy hierarchy but fail learnability constraints?
- Is relevant knowledge encoded in LMs but not causally active in generation?
- What reveals the epistemic limits of language models?
- Can encoder models match human conceptual structure better than larger language models?
- What distinguishes surface generalizations from true linguistic generalizations?
- Why do surface generalizations fail on unusual syntactic structures?
- Why do NLP models fail at recognizing multiple valid interpretations?
- Why do LLMs understand efficient language but fail to produce it?
- Can language models distinguish between novel insight and unjustified conceptual blending?
- Can linear probing detect all the concepts a language model actually uses?
- Why do multimodal models fail on rare and underrepresented concepts?
- Can LLMs use implicit background knowledge the way humans do in ordinary conversation?
- Can smaller open-source LLMs reliably detect agreement across unfamiliar topics?
- Can LLMs learn to signal evaluative commitment through metadiscursive language?
- What structural limits prevent LLMs from abstracting moral principles?
- Can training LLMs to form ad-hoc conventions improve their pragmatic reasoning?
- How do prescriptive ethical constraints differ from descriptive ethical understanding in LLMs?
- How do spoken expert discussions shape what LLMs cannot learn?
- How widespread is task contamination in LLM evaluation benchmarks today?
- Why do benchmark tests fail to detect LLM comprehension gaps?
- Can models identify what information they are missing in underspecified problems?
- Why does homework adherence remain low despite advances in language model capability?
- Can LLMs learn to ask clarifying questions instead of guessing?
- Can large language models actually deliver cognitive behavioral therapy techniques?
- Why do LLMs understand therapy techniques but fail to execute them?
- What training data barriers prevent LLMs from learning real Socratic dialogue?
- Why do reasoning models fail on structurally unfamiliar instances?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- How does instance novelty rather than chain length explain reasoning failure?
- How do dependency errors propagate through incorrectly formalized definitions?
- Why do monological explanations fail to transfer understanding compared to dialogical ones?
- At what complexity does LLM discourse failure become practically harmful?
- What linguistic blind spots do LLMs exhibit in discourse structure?
- Why does semantic decoupling specifically break LLM reasoning abilities?
- Why do LLMs struggle with negation and exception handling?
- How does structural complexity affect LLM performance differently than inferential complexity?
- Can LLMs reliably generate novel working architectures without structured representations?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- Does compressing Walton's schemes into nine categories make LLM classification easier?
- Why do LLM descriptions of argument schemes work better than formal definitions for classification?
- How do single-ordering encyclopedic systems limit ways of knowing?
- Why does entity recognition act as a self-knowledge mechanism in LLMs?
- Can behavioral self-awareness in LLMs extend to recognizing their own contradictions?
- How can we probe LLM representations in channels that training did not target?
- Can representation engineering reliably identify and manipulate self-referential concepts in models?
- Why does AI struggle with wordplay when it has access to word embeddings?
- What happens when technological capacity outpaces ordinary language comprehension?
- Which linguistic abilities are learnable from human-sized data exposure?
- Do rare cultural concepts fail predictably as model scale increases?
- How do we distinguish knowledge encoding from knowledge usage in models?
- What distinguishes conceptual understanding from statistical pattern matching in models?
- How does the knowing-doing gap relate to Potemkin understanding?
- What prevents LLM representations from causally influencing generation outputs?
- When does a model's lack of interpretability become a genuine epistemic problem?
- What role does failure and vulnerability play in real linguistic practice?
- Can understanding language happen entirely within a language system alone?
- What happens when LLMs grade other LLMs in closed evaluation loops?
- When does an LLM judge's error become terrain that an optimizer can map and exploit?
- Why do LLMs swing on minor rewording yet ignore explicit bias correction instructions?
- Why do LLMs strip applicability conditions during memory abstraction?
- Why do language models ignore condensed memory even when it is the only memory?
- How can humans evaluate explanations from systems they did not train?
- Why does FunSearch claim interpretability without measuring human comprehension?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models actually use their encoded knowledge?
Probes can detect that LMs encode facts internally, but do those encoded facts causally influence what the model generates? This explores the gap between knowing and doing.
related mechanism: knowledge can be present without causally influencing behavior; Potemkin extends this to a more observable test (explanation vs. application)
-
Can models pass tests while missing the actual grammar?
Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.
Potemkin understanding adds the recognition-of-failure component that surface generalization accounts don't predict
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
faithful reasoning would prevent Potemkin: the explanation would causally constrain the application
-
Can identical outputs hide broken internal representations?
Can neural networks produce correct outputs while having fundamentally fractured internal structure that prevents generalization and creativity? This challenges our assumptions about what performance benchmarks actually measure.
FER provides the mechanistic cause of Potemkin understanding: the internal representation is fractured across arbitrary subdomains and entangled across unrelated computations, which is why explanation-generation and concept-application are functionally disconnected despite identical surface performance
-
What limits how much models can improve themselves?
Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
the generation-verification gap formalizes why Potemkin understanding paradoxically enables self-improvement: when explanation exceeds application, that gap is a usable training signal — the model's verification ability can supervise its generation ability
-
Why do language models fail to act on their own reasoning?
LLMs produce correct explanations far more often than they produce correct actions. What causes this knowing-doing gap, and can training methods close it?
quantified instance: the 87%/64% gap between correct rationales and correct actions in sequential decision-making is the most precisely measured example of Potemkin understanding; RL fine-tuning narrows the gap, suggesting the facade is partially trainable
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Large Language Model Reasoning Failures
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Word Meanings in Transformer Language Models
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
Original note title
potemkin understanding is a distinct failure mode where correct explanation combined with failed application is incoherent not merely wrong