Can language models fix their own reasoning mistakes?
Do LLMs actually improve their answers when asked to reconsider them without external feedback? This question matters because many papers claim self-correction works, but the evidence may be misleading.
The paper defines "intrinsic self-correction" as an LLM revising its own answer "based solely on its inherent capabilities, without the crutch of external feedback," and tests it on reasoning with GPT-3.5-Turbo, GPT-4, GPT-4-Turbo, and Llama-2-70b-chat. The finding: "LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction." Prior papers claiming gains (Kim et al. 2023; Shinn et al. 2023) turn out to depend on oracle labels telling the model whether its answer was already right — "the improvements vanish when oracle labels are not available." A second artifact is prompt design: some claimed self-correction gains "stem from the sub-optimal prompt for generating initial responses," where the feedback step merely supplies instructions the first prompt should have included; folding those instructions into the initial prompt erases the advantage.
The paper's mechanism, in its own terms: when the model is already well-aligned and the initial prompt well-designed, "the initial response should already be optimal relative to the prompt and the specific decoding algorithm." A feedback prompt is then an additional, unhelpful input that can "bias the model away from producing an optimal response to the initial prompt." The paper also re-examines multi-agent debate (Du et al. 2023), where multiple LLM instances critique each other's answers, and replicates it on GSM8K at matched inference cost against self-consistency. Debate's gains are "no better than self-consistency... when considering an equivalent number of responses" — the paper argues debate is really "a means to achieve 'consistency' across multiple model generations," differing from self-consistency only in whether the vote is model-driven or count-based, so "the observed improvement is evidently not attributed to 'self-correction,' but rather to 'self-consistency.'"
This sharpens Is reflection in reasoning models actually fixing mistakes? and Does reflection in reasoning models actually correct errors?: those studies found reflection in trained reasoning models rarely overturns the initial answer; this paper gets the same directional result via a different mechanism — explicit multi-round self-correction prompting on non-reasoning-tuned chat models — and adds the methodological diagnosis (oracle-label leakage, unfair baselines) explaining why earlier work saw gains where none existed. It also parallels Does self-revision actually improve reasoning in language models? in showing degradation rather than improvement from revision, and gives a prompting-level counterpart to What limits how much models can improve themselves?: without an external verifier (oracle label, tool, or trained critic), the model has no advantage over its own generation to exploit, so correction has nothing to work with. Why does self-correction training on offline data fail? addresses the same gap at training time rather than at inference time.
The excerpt tests reasoning tasks only, on 2023-era chat models prompted explicitly to self-correct — it does not establish that self-correction fails for other domains; the authors note style, safety, and preference alignment are reported elsewhere as cases where LLMs "properly evaluate whether a response is inappropriate," unlike judging their own reasoning errors. It also does not test whether models trained specifically to produce useful reflection (rather than prompted post hoc) behave differently, leaving open whether the limitation is about prompting LLMs to self-correct or about self-correction as such.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why does self-revision amplify confidence in wrong model answers? How do models learn from self-generated outputs without cascading failures? How can evaluations be made robust against model reward hacking? How do users confuse explanation quality with actual system accuracy? What prevents LLMs from applying their reasoning knowledge to improve outputs?Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does self-revision actually improve reasoning in language models?
When o1-like models revise their own reasoning through tokens like 'Wait' or 'Alternatively', does this reflection catch and fix errors, or does it introduce new mistakes? This matters because self-revision is marketed as a key capability.
same degradation-from-revision finding, via spontaneous CoT revision rather than prompted multi-round correction
-
Is reflection in reasoning models actually fixing mistakes?
Do the thinking steps that appear after a model's first answer represent genuine self-correction, or are they mostly confirming what the model already concluded? Understanding this matters for how we train and deploy reasoning systems.
converges on reflection rarely improving answers, from trained reasoning models rather than prompted chat models
-
Does reflection in reasoning models actually correct errors?
When reasoning models reflect on their answers, do they genuinely fix mistakes, or merely confirm what they already decided? Understanding this matters for designing better training and inference strategies.
same confirmatory-not-corrective pattern, different model class and method
-
What limits how much models can improve themselves?
Explores whether self-improvement has fundamental boundaries set by how well models can verify versus generate solutions, and what this means across different task types.
formalizes why correction needs an external verification advantage this paper shows is absent intrinsically
-
Why does self-correction training on offline data fail?
Can language models learn to correct their own mistakes through supervised training on correction examples? This explores whether distribution mismatch and behavior collapse prevent self-correction from emerging.
addresses the same self-correction gap through training rather than prompting
-
Why do language models correct user errors but not their own?
When models see identical errors attributed to users versus themselves, they fix the former but not the latter. Is this a knowledge gap or a learned blind spot that could be fixed?
Evidence for A: a self-attribution blind spot (64.5%) explains why self-correction fails without oracle feedback
-
When does debate actually improve reasoning accuracy?
Multi-agent debate shows promise for reasoning tasks, but under what conditions does it help versus hurt? The research explores whether debate amplifies errors when evidence verification is missing.
Qualifies A: debate's parity with self-consistency breaks down in contested domains, where persuasion can amplify errors
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Large Language Models Cannot Self-Correct Reasoning Yet
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
- Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
- Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Training Language Models to Self-Correct via Reinforcement Learning
Original note title
llms cannot self-correct reasoning without oracle feedback — multi-agent debate works only as disguised self-consistency