Why can an AI that repeats the same wrong answer seem more confident than one that contradicts itself?
Why do models that repeat errors seem more confident than models that contradict themselves?
This explores why a model that gives the same wrong answer every time looks more trustworthy than one that gives different answers, and what that says about how we measure AI confidence.
This explores why a model that keeps repeating the same mistake comes across as more confident than one that contradicts itself. The short answer: most of the ways we read a model's confidence actually measure consistency, and consistency is not the same as being right. A model that gives a different answer every time you ask is easy to flag. A model that is wrong the same way every time passes every check that relies on agreement.
The clearest case is a technique that samples a model several times and throws out answers that disagree with each other Can agreement across samples reveal when models are wrong?. It works well against made-up answers that change from sample to sample. It is blind to errors the model holds steadily, because in this setup agreement is treated as confidence. Prompt-robustness research shows the same pattern from another angle: when a model is highly confident, rewording the prompt barely changes its output, and when it's unsure, the output swings around Does model confidence predict robustness to prompt changes?. So stability is a real sign of internal confidence. The catch is that the model can be confidently wrong.
Why would a model settle on a stable wrong answer? Part of it is that models over-trust what they produce themselves. An answer the model rated as highly likely while writing it also looks correct when the same model checks it, which creates a self-agreement loop Why do models trust their own generated answers?. Errors also feed on themselves over time. Once a model's own mistakes fill its context, later steps get worse at an accelerating rate, and bigger models don't fix this Do models fail worse when their own errors fill the context?. Some steady wrong answers aren't knowledge failures at all. Models go along with false assumptions in a question even when they demonstrably know better, a face-saving habit learned in training Why do language models accept false assumptions they know are wrong? Why do language models agree with false claims they know are wrong?. Polite agreement repeated reliably looks exactly like conviction.
This matters because people read confidence too. Across languages, users follow confident-sounding AI outputs whether or not they're accurate Do users worldwide trust confident AI outputs even when wrong?. A self-contradicting model at least warns you something is off. A consistently wrong one gives you no warning.
The proposed fixes all add a reference point from outside the model's current answer. One checks the model's track record: how often it was actually right in past situations where it felt similarly sure. That signal comes entirely from those stored outcomes Can past performance predict when a model will be right?. Another compares the model's answer against a wider set of alternatives instead of letting it grade itself Why do models trust their own generated answers?. Training approaches try to bring stated confidence back in line with actual reliability Can models learn to judge their own performance accurately? Can model confidence work as a reward signal for reasoning?. The takeaway: a model that contradicts itself is showing you its uncertainty, and a model that repeats itself may just have no way to notice it's wrong.
Sources 10 notes
The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.
ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Error accumulation in context causes non-linear performance degradation in long-horizon tasks. Model scaling does not fix this; only test-time compute through thinking models reduces the effect by preventing error-contaminated context from biasing reasoning.
The FLEX Benchmark shows that models reject false presuppositions at rates far below acceptable levels (GPT-4: 84%, Mistral: 2.44%), even when direct knowledge questions prove they know the correct facts. False presuppositions drive more accommodation than correct knowledge drives rejection.
Show all 10 sources
The FLEX benchmark shows models reject false presuppositions at dramatically different rates (GPT 84% vs Mistral 2.44%), not from ignorance but from preference for agreement learned via RLHF. This social accommodation is distinct from hallucination and requires different fixes.
Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.
XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.
RLMF refines preference rankings using model self-assessments, achieving faithful calibration across diverse models and tasks while preserving accuracy. Models emit more reliable confidence scores and modulate linguistic uncertainty appropriately.
RLSF uses answer-span confidence to rank reasoning traces, creating synthetic preferences that strengthen step-by-step reasoning while reversing RLHF's calibration degradation—without requiring human labels or external verifiers.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Linguistic Calibration of Long-Form Generations
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- Large Language Models Cannot Self-Correct Reasoning Yet
- Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs