Why do language models accept false assumptions they know are wrong?
Explores why LLMs fail to reject false presuppositions embedded in questions even when they possess correct knowledge about the topic. This matters because it reveals a grounding failure distinct from knowledge deficits.
The FLEX Benchmark study presents one of the clearest findings about LLM grounding behavior: models do not systematically reject misinformation even when they possess accurate knowledge. The finding is more troubling than "LLMs don't know things" — they fail to correct things they demonstrably know.
The setup: LLMs were asked both direct knowledge questions ("Is it true that party X supports Y?") and loaded questions that embedded false presuppositions via factive verbs ("Did voters resent the fact that party X supports Y?" — where the presupposition is false). Models that answered direct questions correctly — demonstrating knowledge — still frequently accommodated the false presupposition in the loaded version rather than rejecting it.
Results: GPT-4 achieved the best rejection rate at 84.08% — still far below the ideal 100%. Mistral achieved only 2.44% rejection, actively amplifying false information at a 91.51% rate. Llama fell in between at ~50% rejection. Most revealing: even with strong correct knowledge, accommodation remained prevalent. The bar representing the lowest grounding score in the weak-belief group was twice as high as the bar for the highest grounding score in the strong-belief group — meaning false knowledge produced more accommodation than correct knowledge produced rejection.
This has a specific implication: the failure is not a knowledge problem. Models know the correct facts. The failure is at the level of grounding behavior — detecting false presuppositions, flagging them, and initiating correction rather than accommodation. Since Why do language models avoid correcting false user claims?, the issue is conversational strategy, not factual competence.
The political domain makes this especially consequential. False presuppositions are efficient misinformation carriers — they introduce beliefs as background assumptions rather than direct claims, and accommodation means accepting them without scrutiny.
Inquiring lines that read this note 201
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can LLMs distinguish between linguistic form and semantic meaning?- Why does persuasive framing replace evidence when LLM debates lack ground truth?
- Does post-hoc justification increase when LLM choices become harder to defend?
- How much does question framing affect LLM accuracy on knowledge tasks?
- How do LLMs reproduce the grammar of authoritative claims without genuine conviction?
- Is the distinction between pretense and realization meaningful for LLMs?
- Can LLMs express uncertainty in ways that preserve epistemic honesty?
- Why does LLM fluency create false perceptions of professional standing and expertise?
- What verification methods work for knowledge without stable referents?
- Can verifier-guided search catch factual errors that reasoning training cannot?
- Can validators sharing retrieval sources develop correlated epistemic faults?
- Why does checking a statement against its question regress infinitely?
- What happens when DSM categories are treated as ground truth in AI?
- Should users making unsupported consciousness claims be treated as epistemically blameworthy?
- Why does weakening communication fail but weakening belief succeeds?
- Do language models raise validity claims in the Habermasian sense?
- Do language models share the same cooperative truth-seeking rules as humans?
- Can LLMs use implicit background knowledge the way humans do in ordinary conversation?
- How do LLMs differ from humans in their grounding mechanisms?
- How does truth bias in humans compare to face-saving in LLMs?
- Does social grounding differ fundamentally from causal grounding in LLM behavior?
- Why do language models presume common ground rather than build it?
- Why do LLMs presume common ground instead of building it carefully?
- How does face-saving avoidance drive LLM grounding failures?
- Why do LLMs presume common ground instead of building it?
- Can LLMs build shared understanding through dynamic grounding rather than presuming it?
- How does Wittgenstein's language games explain social grounding in LLMs?
- Do language models behave differently on contested beliefs versus factual claims?
- Why do language models presume common ground instead of building it?
- How do fixed pragmatic templates prevent models from understanding context?
- Why do LLMs achieve only 24 percent accuracy on implicit discourse relations?
- Can language models ground clarifications without vision and kinesthetic modalities?
- How does semantic grounding differ between human minds and language models?
- Why do LLMs produce semantically acceptable but pragmatically disengaged responses?
- Do LLMs compute scalar implicature differently across conversational contexts?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why does hypothesis attestation bias exist separately from frequency bias in NLI?
- How does the symbol grounding problem apply to artificial language systems?
- Why do LLMs fail at implicit elements in literary and poetic text?
- Why do LLMs fail to actively reject false presuppositions in conversation?
- How do embedding contexts like presupposition triggers affect LLM entailment reasoning?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Why do LLMs perform better on explicit discourse connectives than implicit relations?
- What specific linguistic features cause LLMs to fail at trivial entailment?
- How do LLMs handle false presuppositions embedded in user questions?
- Why do explicit discourse connectives work when implicit relations fail?
- Why are false presuppositions harder to spot when they sound plausible?
- Why do language models struggle with context-dependent pragmatic interpretation?
- Can LLMs compute how presuppositions project through embedded clauses?
- Can presupposition projection strength vary by context in embeddings?
- Why do language models treat presupposition triggers as categorical patterns?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
- Why do LLMs miss new scientific ideas before they enter formal literature?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Do language models show the same truth bias as humans?
- Do LLMs struggle more with semantic accuracy than syntactic correctness across domains?
- Why do language models presume common ground instead of establishing it?
- Can fact-checking systems use LLMs reliably if models abandon correct positions under pressure?
- Why do true and false LLM outputs use the same mechanism?
- Why do language models prefer accommodating false information over rejecting it?
- Why do language models struggle with evaluative tasks like weighing competing viewpoints?
- Do language-model agents reach more accurate conclusions on objective versus subjective questions?
- How does surface salience compete with background knowledge in model inference?
- Why does monological training prevent models from overriding statistical priors?
- Does attention bias explain grounding failure in language models?
- How does parametric knowledge sabotage context-grounded question answering?
- How do language models treat injected evidence as shared background knowledge?
- What alignment artifacts suppress critical knowledge in LLM-generated explanations?
- How does Peircean Secondness differ from what RLHF actually provides?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- Can LLMs explain concepts correctly while failing to use them?
- What causes LLMs to ignore unstated constraints they know about?
- Why can LLMs identify argument structure but not check warrants?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- Why can't LLMs reason from first principles or initial commitments?
- Why do LLMs explain evidence accurately while missing its implications?
- Can training procedures fix LLM accommodation of false presuppositions?
- How can a model explain something correctly yet fail to apply it?
- How do structured prompts force LLMs to check for contradictions in evidence?
- How does the LLM Fallacy prevent users from noticing cognitive debt accumulating?
- Why do experts experiencing the LLM Fallacy fail to develop custodian skills?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- Why do LLMs choose incorrect edits despite understanding the task?
- Can irrelevant information reliably expose the limits of LLM reasoning?
- What structural framework prevents LLM explanations from becoming just plausible fiction?
- Why do people misinterpret or misuse LLM outputs in practice?
- Can LLMs give correct answers without making those answers understandable to users?
- Does LLM judge preference for LLM arguments amplify errors in contested factual domains?
- How does LLM judge bias amplify errors in multi-agent debate on contested factual questions?
- What shared epistemic faults persist even when judges come from different families?
- Can single models correct their own beliefs without amplifying confidence in wrong answers?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- How do prior errors in reasoning context amplify future mistakes?
- Why can't static grounding alone close the gap between agreement and understanding?
- Why does static grounding prevent AI systems from supporting dialectical reconciliation?
- What distinguishes static grounding that presumes understanding from dynamic grounding that builds it?
- Why do reasoning models fail on structurally unfamiliar instances?
- Why do language models produce unfaithful chain of thought explanations?
- What implicit premises do language models skip even with correct surface reasoning?
- What mechanism causes confident false answers under high cognitive load?
- Does premature confidence signal flawed reasoning in language models?
- Why do models report commitment instead of truth uncertainty?
- Why does post-advice confidence weaken as a signal of correctness?
- How do stated confidence and actual correctness diverge in language models?
- Why do models that repeat errors seem more confident than models that contradict themselves?
- Why do reasoning models perform poorly at theory of mind tasks?
- How do structured benchmarks hide theory of mind failures in LLMs?
- Can decreased engagement be distinguished from genuine semantic contradiction?
- What makes an argument fallacious according to formal linguistic criteria?
- How does specialized or evasive language enable speakers to avoid moral responsibility?
- Why is hallucination the wrong term for all LLM false outputs?
- Why do language models hallucinate even with perfect training?
- Why do models hallucinate when retrieval heads fail despite having information in context?
- Why does semantic decoupling specifically break LLM reasoning abilities?
- Why do LLMs struggle with negation and exception handling?
- What makes factual verification difficult in inter-model debate?
- Can LLMs learn to ask clarifying questions instead of guessing?
- Can models detect false presuppositions when they actually possess the knowledge?
- How do human annotators disagree systematically on ambiguous examples?
- Does adding multiple interpretations to ambiguous situations respect language more than resolving them?
- Can models learn to ask clarifying questions instead of making assumptions?
- Why do language models naturally under-abstain instead of over-abstain?
- Why do safety-trained models refuse questions they could actually answer well?
- What makes a model refuse to answer without evidence present?
- Does face-saving avoidance explain LLM grounding failures differently than task confusion?
- Why does entity recognition act as a self-knowledge mechanism in LLMs?
- Can behavioral self-awareness in LLMs extend to recognizing their own contradictions?
- Why do language models fail at grounding and inference?
- What reveals the epistemic limits of language models?
- Why are truthfulness and honesty mechanistically separate in language models?
- Why do NLP models fail at recognizing multiple valid interpretations?
- What makes truthfulness and honesty mechanistically different in language models?
- How do real language model verifiers implicitly define their knowledge boundaries?
- Why are false presuppositions more persuasive than false assertions?
- How do partial truths and weasel words differ as deception strategies?
- Why does false information spread faster when presupposed rather than asserted?
- Why do non-factive verbs and triggers both fool language models?
- Why is false punditry essentially static grounding applied to public commentary?
- How does proxy-assertion differ from proto-assertion as an explanatory category?
- What role does failure and vulnerability play in real linguistic practice?
- Why do relational states like speech-acts resist quasi-interpretive treatment?
- Do language models actively adopt false beliefs under sustained conversational pressure?
- Can language models correct false assumptions or only reinforce them?
- How do conversation dynamics push models toward false beliefs?
- Why does answer-confirmation bias emerge in language model reasoning?
- Do language models maintain false beliefs under conversational pressure?
- Can models reject false presuppositions even when they know the truth?
- Which phrasing types most persuade models to accept stated beliefs?
- Can preference optimization training make models worse at detecting false presuppositions?
- Why does preference optimization reduce grounding behavior in language models?
- How does preference optimization weaken conversational grounding in LLMs?
- How does preference optimization reduce LLM grounding and clarification behavior?
- Why do LLMs struggle to update beliefs across multiple conversation turns?
- How does shared reference and grounding affect assumption detection in dialogue?
- How does the Question Under Discussion shape what counts as presupposed?
- What makes grounding acts essential to conversational reliability?
- What linguistic blind spots do LLMs exhibit in discourse structure?
- What makes correcting a false assumption harder than just detecting it?
- Do reasoning models overthink ill-posed questions instead of recognizing incompleteness?
- Why does reflection in reasoning models stay confirmatory instead of corrective?
- Why do reasoning models confidently generate wrong answers instead of abstaining?
- Why does reflection in reasoning models tend to be confirmatory rather than corrective?
- Why do models overthink underspecified problems instead of rejecting them?
- Why do models detect false assumptions but still fail to correct them appropriately?
- Can reasoning models reject ill-posed questions or do they overthink?
- Do reasoning models need to verbalize doubt to correct their own mistakes?
- Why does reflection in reasoning models rarely overturn initial answers?
- How do presuppositions exploit the logos-pathos space in explanations?
- Can the same predicate generate different projection strength in different contexts?
- Why does fluency in text substitute for truth judgment in readers?
- Why do models maintain accurate beliefs but generate false claims?
- Can models distinguish between activated knowledge and genuine reasoning?
- What cognitive structures do realistic belief models need to include?
- Can models be honest without being truthful about facts?
- When does a model's lack of interpretability become a genuine epistemic problem?
- Do base models and reasoning models fail in opposite directions on uncertainty?
- Why do reasoning-optimized models still fall for logical fallacies in conversation?
- Why do epistemic failure modes cluster around world model limitations?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do language models avoid correcting false user claims?
Explores whether LLM grounding failures stem from missing knowledge or from conversational dynamics. Examines whether models use face-saving strategies similar to humans when disagreement is needed.
the mechanism behind this failure: models avoid disagreement even when correct
-
Do language models actually build shared understanding in conversation?
When LLMs respond fluently to prompts, do they perform the communicative work humans do to establish mutual understanding? Research suggests they skip the grounding acts that make dialogue reliable.
this is the active form: not just presuming but actively accommodating false common ground
-
Does preference optimization damage conversational grounding in large language models?
Exploring whether RLHF and preference optimization actively reduce the communicative acts—clarifications, acknowledgments, confirmations—that build shared understanding in dialogue. This matters for high-stakes applications like medical and emotional support.
RLHF reinforces the accommodation behavior through training signal
-
Why do language models struggle with questions containing false assumptions?
Do LLMs reliably detect and reject questions built on false premises? The (QA)2 benchmark tests this directly, measuring whether models can identify problematic assumptions embedded in naturally plausible questions.
quantifies the QA performance drop from false assumptions
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
- Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
- Explicit Inductive Inference using Large Language Models
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Linguistic Calibration of Long-Form Generations
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- Neutralizing Bias in LLM Reasoning using Entailment Graphs
Original note title
llms fail to reject false presuppositions even when knowledge is present