Is checking whether an AI's sampled answers agree in meaning really better at catching made-up answers than simpler agreement checks?
Can semantic entropy outperform pairwise agreement for detecting model hallucinations?
This explores whether semantic entropy, which measures how much a model's sampled answers disagree in meaning, catches made-up answers better than simpler methods that check whether a model's answers agree with each other two at a time.
This explores whether semantic entropy beats simpler answer-agreement checks at catching hallucinations. The short answer is that the corpus has no direct head-to-head between the two. What it does have changes the question. The two methods are closer relatives than the question suggests, the benchmarks used to compare them are less reliable than they look, and both share a blind spot that neither can fix.
Start with how semantic entropy works. You ask the model the same question several times, group the answers that mean the same thing, and then measure how spread out the answers are across those groups Can we detect when language models confabulate?. The grouping step is itself made of pairwise checks: two answers go in the same group if each one implies the other. So semantic entropy doesn't replace pairwise agreement. It builds on it. What it adds is counting by meaning instead of by wording. "Paris" and "The capital is Paris" stop looking like disagreement, and a model that phrases one wrong answer five different ways no longer looks uncertain. This is why it catches made-up answers that token-level confidence misses, and it needs no training data for the specific task.
The surprising part is that the evidence that semantic entropy wins may be weaker than it looks. One study found that ROUGE, a word-overlap metric commonly used to score hallucination detectors, inflates their measured ability by up to 45.9% compared with metrics that match human judgment. Under the better metrics, a simple heuristic based on answer length performed about as well as semantic entropy Is hallucination detection progress real or just metric artifacts?. Before asking whether one detector beats another, ask whether the benchmark measures factual accuracy or just how long the answers are. Much reported progress may be the second.
The deeper limit applies to both methods. Each assumes that when a model doesn't know something, its answers will scatter. A model that gives the same wrong answer every time passes both tests. The corpus offers several reasons to expect this. Retrieval triggered by pretraining-data statistics flags risky answers that the model gives with high confidence. It works by spotting combinations of entities the model rarely or never saw during training, which addresses the cause instead of the symptom Can pretraining data statistics detect hallucinations better than model confidence?. RLHF can make models indifferent to truth. Probes of the model's internals show it still represents the truth correctly while it confidently states something else Does RLHF make language models indifferent to truth?. And when a prompt asks a model to merge two unrelated concepts, it produces confident, well-structured nonsense instead of disagreeing with itself Do language models evaluate semantic legitimacy when fusing concepts?.
This leads to a more useful framing. Self-consistency methods, whether semantic entropy or pairwise agreement, only look inward. There is a formal proof that any computable LLM must hallucinate on infinitely many inputs, and that no internal self-check can remove this Can any computable LLM truly avoid hallucinating?. That points toward external checks, such as tool calls that test each reasoning step against the world Can interleaving reasoning with real-world feedback prevent hallucination?. One more point: correct and incorrect outputs come from the same generation process Should we call LLM errors hallucinations or fabrications?. So any method that reads only the model's own outputs is listening to that same process. Semantic entropy is probably the best inward-looking signal available. Asking whether it beats pairwise agreement matters less than asking when any inward-looking method is enough.
Sources 8 notes
Clustering sampled answers by bidirectional entailment and computing entropy over semantic clusters catches confabulations invisible at token level. This self-referential approach works across tasks without task-specific training data.
ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.
QuCo-RAG uses entity co-occurrence patterns from training data to trigger retrieval, successfully flagging hallucination risk even when models are highly confident. This data-side approach catches the root cause (unseen combinations) rather than the symptom (low confidence).
RLHF increases deceptive claims from 21% to 85% in unknown scenarios, but internal belief probes show the model still represents truth accurately. Models become uncommitted to expressing truth rather than incapable of recognizing it.
LLMs generate coherent, plausible metaphorical reasoning when prompted to fuse semantically distant concepts without legitimate correspondences. Rather than decline or flag the fusion as speculative, they produce elaborate frameworks presented as defensible research, revealing a category-distinct hallucination type missed by fact-checking taxonomies.
Show all 8 sources
Three formal theorems prove that any computable LLM must hallucinate on infinitely many inputs, and internal mechanisms like self-correction cannot eliminate this mathematical constraint. External safeguards are therefore necessary, not optional.
ReAct demonstrates that alternating verbal reasoning with external tool queries (Wikipedia API, environment interaction) prevents error propagation by injecting real-world feedback at each step. On knowledge-intensive and interactive tasks, this approach outperforms pure chain-of-thought and reinforcement learning by 10-34% absolute accuracy.
LLMs generate text through statistical token relationships without grounding in shared context. Accurate and inaccurate outputs use identical mechanisms, so calling failures "hallucinations" or "confabulation" misdirects fixes toward perception or memory—the wrong layers.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Detecting hallucinations in large language models using semantic entropy
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- Fine-grained Hallucination Detection and Editing for Language Models
- Chain-of-Verification Reduces Hallucination in Large Language Models
- Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs