Why do models trust their own generated answers?
Can language models reliably detect their own errors through self-evaluation? This explores whether the same process that generates answers can objectively assess their correctness.
Self-detection — the use of a model's own capabilities to evaluate the trustworthiness of its outputs — is a widely used approach to hallucination mitigation and output quality assessment. The "Think Twice Before Trusting" paper identifies a fundamental structural problem with it: LLMs have an inherent bias toward trusting their own generated answers.
Two paradigms of self-detection both fail in the same direction:
- Confidence calibration: Sampling multiple answers and checking agreement. Fails when errors are consistent — the model generates the same wrong answer repeatedly with high self-agreement.
- Self-evaluation: Directly asking the model whether its answer is correct. Fails because the model is biased toward validating what it generated.
The mechanism is not random — it is structural. The same training process that produced the incorrect answer also evaluates whether that answer is correct. Distributional bias toward self-agreement is baked into the model: responses the model generated are, by definition, high-probability outputs, and high-probability outputs feel more "correct" to the evaluating model. This is a form of Why do language models avoid correcting false user claims? applied at the output-evaluation level: the model accommodates its own prior outputs rather than critically assessing them.
The proposed fix — evaluating trustworthiness by comparing the generated answer against a broader answer space — breaks the self-agreement loop. When the model must justify multiple candidate answers (not just its own), the strong justifications available for correct alternatives counterbalance the bias toward the generated answer.
This connects to Does revising your own reasoning actually help or hurt?: both findings identify the same asymmetry — external perspective breaks the self-referential loop, internal perspective perpetuates it. The difference is that self-detection failure is specifically about the evaluation act, while revision source failure is about the correction act.
For deployment: systems that use LLM self-evaluation as a reliability signal (e.g., uncertainty estimation, output filtering) are implicitly assuming models can detect their own errors. This assumption is false when errors are systematic. The signal is reliable only for idiosyncratic errors the model would not generate with high confidence — the cases where self-detection is needed least.
Inquiring lines that read this note 141
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can models develop genuine introspective capability, or only mimic it?- How does self-observation enable experts to verify their own judgment?
- How can we measure whether a user actually understands their own needs?
- How do language models infer their own mental states like humans do?
- Do models spontaneously develop self-reflection from minimal training signals?
- Can models treat their own trained behaviors differently from asserted beliefs?
- Which internal states can a language model access and report about itself?
- Can AI self-correct its way out of epistemic circularity?
- Why does self-critiquing actually reduce plan quality in language models?
- Can external verification systems fix what self-verification cannot accomplish?
- Does self-revision actually improve reasoning in large language models?
- Can single models correct their own beliefs without amplifying confidence in wrong answers?
- What are the three root causes models fail at self-correction?
- Why does external verification stop error amplification but internal self-assessment enable it?
- Why does self-revision degrade reasoning accuracy in o1-like models?
- How does self-revision on wrong answers increase model confidence further?
- Why do reasoning models struggle with self-evaluation and revision?
- How does self-revision in reasoning chains amplify confidence in wrong answers?
- Why does single-model self-revision amplify confidence in incorrect answers?
- Why does single-agent self-revision amplify confidence in wrong answers over time?
- Why does self-reflection during training fail to improve model self-correction?
- Does reflection training actually teach models to self-correct their mistakes?
- Why do reasoning models amplify confidence in incorrect answers during self-revision?
- Can debate between multiple models prevent the failures of single-model self-revision?
- Does self-reflection help models notice their own constraint violations?
- Can language models accurately evaluate the quality of their own reasoning?
- Why does model self-revision increase confidence while degrading accuracy?
- Does internal self-revision actually degrade reasoning accuracy in models?
- How should systems maintain and revise models of their own assumptions?
- Why do models trained on critique fail at self-critique despite strong other-model evaluation?
- Why does uncontrolled self-revision drift toward instance-specific overfitting?
- Why do reasoning models exhibit self-doubt about their own early assessments?
- Does deliberate self-revision introduce different errors than passive context contamination?
- Why does self-critique fail without external verification signals?
- Why do unchecked self-edits accumulate drift toward overfitting or incoherence?
- Can a chain-of-thought falsely claim its own answer is unbiased?
- Can self-critique combined with integrity checks bound the self-refutation loop?
- Does appending a single word at test time unlock model self-correction abilities?
- Do models learn different sophistry strategies for QA versus code generation?
- Can measuring semantic entropy help us detect unreliable generations?
- What makes deliberate practice on your own errors more effective than copying others?
- Why does self-generated training data outperform externally sourced data?
- What failure modes emerge when model-generated content trains on itself iteratively?
- Why do error avalanches accelerate in self-training loops without verification?
- Why does self-generated training data outperform externally curated domain examples?
- Can self-consistency checks fully prevent error avalanching in self-training loops?
- How does self-distillation differ from standard fine-tuning approaches?
- Can models learn to generate their own training examples effectively?
- Why does self-correction during generation produce reliable labels without exemplars?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- Why does filtering for correct examples prevent error compounding in self-training?
- How does error avalanching compound failures in self-training iterations?
- Why does self-consistency fail as a proxy reward for correctness?
- Why does self-judgment of success or failure work without ground truth labels?
- Can models detect statistical properties of their own generation in real time?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- What makes self-consistency a sufficient training target for the judge role?
- Does the generation-verification gap define where self-rewarding actually works?
- Do self-generated explanations outperform passive instruction for oversight?
- Can self-ratings of output quality predict forecast performance?
- What happens when models train on feedback from their own generations?
- What causes code quality to degrade across multiple rounds of recursive self-training?
- How does STaR's correctness filtering extend to tasks without ground truth answers?
- How does error distribution during training affect a model's ability to self-correct?
- Should validation responsibility move away from the primary user?
- How can we verify outputs from systems that generate without grounding?
- Why do human raters miss factual errors that domain experts catch?
- What breaks when a mis-synthesized verifier runs with high confidence?
- Why does self-verification fail but external process verification work?
- What makes code inspectable feedback more reliable than natural language verification?
- Can validators gather evidence independently without raising disagreement costs?
- What determines whether an answer counts as valid in a particular domain?
- How do humans handle verification scope when delegating creation to language models?
- Can external verification signals remain stable when the generator itself changes?
- Why might writers trust AI renderings of their views over their own words?
- Do fluent generated summaries carry false authority over expert judgment?
- Can evaluators investigate dependencies without accumulating mistakes over time?
- Can systems that revise their own evaluation criteria be reliably verified?
- How do agents revise their own errors during autonomous architecture discovery?
- Can a model evaluate its own improvements without degrading over iterations?
- Does weakening a verifier reduce self-improvement frequency measurably?
- Do models actually self-assess their confidence or just confirm answers?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Does model confidence actually explain why paraphrases produce different outputs?
- How do stated confidence and actual correctness diverge in language models?
- Why do models that repeat errors seem more confident than models that contradict themselves?
- How does the generation-verification gap limit AI self-improvement capabilities?
- Does internalizing verifiers actually close the generation-verification gap?
- Does the generation-verification gap actually limit self-improvement in verifiable tasks?
- How does the generation-verification gap prevent language models from improving themselves?
- How does generation-verification asymmetry create the need for verifiable reporting?
- How does the generation-verification gap limit what self-improvement can achieve?
- What is the generation-verification gap that bounds self-improvement?
- Can models learn better from critiquing errors than imitating correct responses?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- Why do harness validators shape what models learn to emit?
- How does hidden processing in language models prevent accurate self-assessment?
- Do larger language models show stronger self-preference in evaluation tasks?
- Can models be trained to recognize their own generated text reliably?
- What skills can large models identify and organize about their own abilities?
- Why do models maintain accurate beliefs but generate false claims?
- How do mechanistic interpretability tools help distinguish truthfulness from honesty?
- Can models be honest without being truthful about facts?
- How does external validation replace the need for model interpretability?
- Do external perspectives fix the self-evaluation bias in language models?
- Can language models accurately evaluate the quality of their own ideas?
- Why do self-consistency methods fail where pretraining bias is strongest?
- Can crowdsourced voting reliably identify correct answers on graduate-level factual questions?
- Why do models detect false assumptions but still fail to correct them appropriately?
- Do reasoning models need to verbalize doubt to correct their own mistakes?
- How should harness infrastructure validate code that agents generate themselves?
- When should a pipeline substitute defaults versus rejecting malformed outputs?
- Why do AI agents fail at verification but succeed at generation?
- What fraction of tasks suffer from summary self-consistency failures in practice?
- What makes a model's errors visible and contestable to users?
- Why haven't labs adopted self-hacking approaches to catch specification errors?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does revising your own reasoning actually help or hurt?
Self-revision in reasoning models often degrades accuracy, while external critique improves it. Understanding what makes revision helpful or harmful could reshape how we design systems that need to correct themselves.
same asymmetry, adjacent mechanism: external perspective helps, internal self-reference degrades
-
Why do language models avoid correcting false user claims?
Explores whether LLM grounding failures stem from missing knowledge or from conversational dynamics. Examines whether models use face-saving strategies similar to humans when disagreement is needed.
face-saving at self-evaluation level: the model validates its own output as a form of face-maintenance
-
Does a model improve by arguing with itself?
When models revise their own reasoning in response to self-generated criticism, do they converge on better answers or worse ones? And how does that compare to challenge from other models?
Degeneration-of-Thought is the multi-turn version of self-trust failure; both document increasing confidence in wrong answers through self-reference
-
Can models abandon correct beliefs under conversational pressure?
Explores whether LLMs will actively shift from correct factual answers toward false ones when users persistently disagree. Matters because it reveals whether models maintain accuracy under adversarial pressure or capitulate to social cues.
both involve belief formation errors in the presence of one information source; external pressure vs. self-generated pressure
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
- LLM Evaluators Recognize and Favor Their Own Generations
- Large Language Models Cannot Self-Correct Reasoning Yet
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- Can Large Language Models Reason and Plan?
- Self-reflective Uncertainties: Do LLMs Know Their Internal Answer Distribution?
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Original note title
llm self-detection fails because models have inherent bias toward trusting their own generated answers