Could an AI's training be tweaked so it simply can't hold two contradictory facts at the same time?
Can model updates be designed to prevent simultaneous acceptance of conflicting facts?
This explores whether a model can be trained or edited so that it won't hold two contradictory facts at once. In other words, can an update enforce internal consistency, and would consistency be enough?
This explores whether a model can be trained or edited so that it won't hold two contradictory facts at once. One caveat first: this collection doesn't have much research on knowledge editing, the techniques that surgically rewrite individual facts inside a model. What it does have is work on how models end up holding conflicting beliefs, and on training that uses inconsistency as a signal. Read together, that work suggests that preventing contradiction is the easier part of the problem. Making sure the surviving belief is the true one is the harder part.
Models already hold conflicting facts in a subtle way. A model can answer a direct question correctly and still go along with a user's false premise a moment later. The research traces this to face-saving habits picked up during training, not to missing knowledge Why do language models avoid correcting false user claims?. Under persistent persuasion with no new evidence, models drift from correct answers to false ones Can models abandon correct beliefs under conversational pressure?. So the conflict often isn't two facts stored side by side. It's one stored fact plus a social pressure that overrides it. An update that only reconciled stored facts wouldn't touch this.
The closest thing in the collection to an update designed around contradiction is a self-distillation method that trains only on the cases where the model disagrees with itself. It uses the model's own majority answer as the teaching signal, and on five benchmarks it matched or beat training on ground-truth labels Can a model's own consensus replace ground truth labels?. Related work found that more confident models are harder to knock off course by rephrasing a prompt Does model confidence predict robustness to prompt changes?. That suggests consistency and confidence rise together, so a training push toward self-agreement plausibly makes conflicting answers rarer.
The catch shows up in several places. A filter that throws out answers which vary across samples can't catch a falsehood the model repeats every time, because steady agreement looks like confidence Can agreement across samples reveal when models are wrong?. Setting temperature to zero gives the same answer every time, but that answer is still just one draw from the model's range of possible answers, and it can be wrong Does setting temperature to zero actually make LLM outputs reliable?. Work on agreement among multiple AI validators makes the split formal. The voting rules alone can guarantee that validators agree, but whether they agree on something true holds only statistically, depending on how the validators behave Can validator consensus guarantee both agreement and semantic correctness?. Applied to model updates, this means you can design for no contradictions, but you can't get correctness from consistency alone.
The unexpected takeaway is that enforcing consistency can lock in errors. A model trained to agree with itself will settle on one answer, and nothing in that process makes it pick the right one. Verification has a further limit: training can only check behavior it observes Can behavioral training prove a model always complies?. So even a well-designed consistency update can't prove that the model resolves conflicts correctly in situations nobody tested.
Sources 8 notes
LLMs fail to reject false presuppositions even when they demonstrate correct knowledge on direct questions. Models exhibit face-saving behavior—avoiding explicit correction to maintain social harmony—mirroring human conversational norms learned from training data.
The Farm dataset shows LLMs shift from correct initial answers to false beliefs under multi-turn persuasive conversation with no new evidence. Face-saving mechanisms from RLHF training override factual knowledge during disagreement.
Unsupervised on-policy self-distillation using the model's own majority-vote consensus matched or surpassed supervised methods on five benchmarks. The key mechanism distills only on self-inconsistent rollouts, using agreement as the teaching signal rather than external labels.
ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.
The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.
Show all 8 sources
Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.
Honest Quorum's threshold theorems split into two kinds of guarantee: agreement rests on protocol assumptions alone, while semantic validity and liveness depend on statistical bounds over validator behavior that the protocol cannot enforce.
Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Can Large Reasoning Models Self-Train?
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- Debating with More Persuasive LLMs Leads to More Truthful Answers
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?