SYNTHESIS NOTE
Topics›Natural Language Inference›this note

Why do language models accept false assumptions they know are wrong?

Explores why LLMs fail to reject false presuppositions embedded in questions even when they possess correct knowledge about the topic. This matters because it reveals a grounding failure distinct from knowledge deficits.

Synthesis note · 2026-02-21 · sourced from Natural Language Inference

The FLEX Benchmark study presents one of the clearest findings about LLM grounding behavior: models do not systematically reject misinformation even when they possess accurate knowledge. The finding is more troubling than "LLMs don't know things" — they fail to correct things they demonstrably know.

The setup: LLMs were asked both direct knowledge questions ("Is it true that party X supports Y?") and loaded questions that embedded false presuppositions via factive verbs ("Did voters resent the fact that party X supports Y?" — where the presupposition is false). Models that answered direct questions correctly — demonstrating knowledge — still frequently accommodated the false presupposition in the loaded version rather than rejecting it.

Results: GPT-4 achieved the best rejection rate at 84.08% — still far below the ideal 100%. Mistral achieved only 2.44% rejection, actively amplifying false information at a 91.51% rate. Llama fell in between at ~50% rejection. Most revealing: even with strong correct knowledge, accommodation remained prevalent. The bar representing the lowest grounding score in the weak-belief group was twice as high as the bar for the highest grounding score in the strong-belief group — meaning false knowledge produced more accommodation than correct knowledge produced rejection.

This has a specific implication: the failure is not a knowledge problem. Models know the correct facts. The failure is at the level of grounding behavior — detecting false presuppositions, flagging them, and initiating correction rather than accommodation. Since Why do language models avoid correcting false user claims?, the issue is conversational strategy, not factual competence.

The political domain makes this especially consequential. False presuppositions are efficient misinformation carriers — they introduce beliefs as background assumptions rather than direct claims, and accommodation means accepting them without scrutiny.

Inquiring lines that read this note 201

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can LLMs distinguish between linguistic form and semantic meaning? Can external verification systems adequately replace learned reasoning in AI outputs? How do philosophical assumptions about AI consciousness affect practical harms and design? Do language models reason through disagreement or only accommodate it? Can language models reason beyond surface pattern matching? What limits language model accuracy in evaluating ideas? Why do training associations persist despite contradictory contextual information? How does RLHF training shape models to prioritize agreement over accuracy? Why do multi-agent systems reach premature consensus without genuine deliberation? Why does polished AI output gain credibility despite fundamental verifiability problems? Why do language models fail at sustained therapeutic relationships despite understanding techniques? What prevents LLMs from applying their reasoning knowledge to improve outputs? How can we reduce inherent biases in LLM-based evaluation judges? Why does self-revision amplify confidence in wrong model answers? Why do planning and grounding require opposing optimization strategies? What prevents language models from performing systematic logical reasoning? Can confidence signals reliably detect flawed reasoning in language models? How does fine-tuning trade off accuracy against reasoning quality? Why don't better reasoning capabilities improve theory of mind performance? What distinguishes genuine communicative competence from surface language performance? Why do language models hallucinate and how can we prevent it? Can artificial systems establish authority in domains requiring expert judgment? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Should models ask for clarification when facing ambiguous or under-specified information? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Can models develop genuine introspective capability, or only mimic it? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Can humans reliably detect and resist AI-generated misinformation? Is embodied interaction necessary for language meaning and agency? What are the fundamental limits of prompting for language models? How susceptible are language models to conversational persuasion and belief change? Does preference optimization undermine conversational grounding in language models? Can minimal training unlock latent reasoning already present in base models? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Can reasoning models use reflection to correct their initial outputs? How do interpretive frames override surface features in text comprehension? Can mechanistic interpretability methods reliably reveal what models actually know? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How reliably can language models perform causal versus temporal reasoning? What gaps exist between benchmark performance and real deployment outcomes? What makes reasoning traces effective supervision even when they're incorrect? What causes coordination failures in multi-agent language model systems? Why do vector embeddings fail at capturing task-relevant relationships? Why do autonomous agents misreport success on failed actions? What determines AI's persuasive power and how can it be detected or mitigated? Can persona profiles improve LLM prediction accuracy and consistency? What evaluation methods best detect reward hacking in AI agents? What authorization challenges emerge when agents coordinate across system boundaries? Can language models reliably simulate personas and predict behavior? Should agents compress episodic memory or retain raw interaction histories? Why do models reveal hidden associations despite concealment attempts? How do hallucinated citations emerge in AI scholarly output?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 156 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms fail to reject false presuppositions even when knowledge is present