If an AI's correction examples come from another model, not its own mistakes, does practicing on them actually teach it to self-correct?
Does training on model-generated correction traces actually work?
This asks whether a model can learn to fix its own mistakes by being fine-tuned on examples of corrections, where those examples were written by a model rather than produced by the model during live practice, and what the collection says about why that does or doesn't work.
This asks whether a model can learn to fix its own mistakes from a fixed set of model-written correction examples, rather than by practicing on mistakes it makes as it trains. The short answer from the collection is: mostly no, at least not with ordinary supervised fine-tuning. The reason is more interesting than "the data is bad." Even when the correction examples look good, the mistakes in that fixed dataset aren't the mistakes the model will make once it's deployed. The model learns to fix someone else's errors, and it tends to fall into one habitual way of correcting instead of responding to what actually went wrong. What did work was online reinforcement learning over several turns. There, the model makes its own mistakes and is rewarded for actually fixing them Why does self-correction training on offline data fail?. Self-correction looks less like a skill you can copy from examples and more like one you have to practice.
That sits inside a bigger surprise in the collection: training traces may not teach what they appear to teach. Models trained on deliberately corrupted or irrelevant reasoning traces do about as well as models trained on correct ones, and sometimes generalize better. That suggests the traces act more like scaffolding for computation than like lessons in reasoning Do reasoning traces need to be semantically correct?. A related line of work finds that the steps in a reasoning trace don't reliably cause the final answer. Invalid traces often still end in correct answers, which points to learned formatting rather than working logic Do reasoning traces actually cause correct answers?. If so, a correction trace may mostly teach the model to look like it is correcting itself, which helps explain why offline correction data collapses into one fixed style.
A correct trace isn't automatically good training material either. One study found that traces which keep exploring after the answer is already settled make fine-tuning worse, even though everything in them is correct. Cutting just that unnecessary tail helped more than cutting a random chunk of the same length Does every correct chain-of-thought trace improve fine-tuning?. So what shapes learning is the structure of a trace, not only whether it is right. Choosing which traces to keep matters too: checking the model's confidence step by step catches reasoning breakdowns that an average over the whole trace hides Does step-level confidence outperform global averaging for trace filtering?.
Learning from mistakes still works in some forms. One method has the model deliberately make errors on a handful of examples in its prompt, reflect on them, and write down general principles. That improves reasoning with no fine-tuning at all Does learning from mistakes improve in-context learning?. What changes is the target: the model distills a lesson from its own errors instead of copying someone else's fixes. Self-generated training data also works when something trustworthy checks it. LLM judges given reference answers, for example, can supervise self-improvement about as well as a trained reward model Can reference examples make LLM judges reliable enough for self-improvement?.
The takeaway you might not have expected: whether correction training works depends less on how good the examples are and more on whose mistakes they are. Teaching a model to fix errors works best when the errors are its own, encountered while it practices. Online training has its own failure modes, though. Problems that are nearly impossible can reward lucky shortcuts instead of real reasoning Do overly hard RLVR samples actually harm model capabilities?. And reinforcement learning tends to narrow a model toward one dominant output style Does RL training collapse format diversity in pretrained models?, so collapsing into a single correction style isn't a problem only offline training has.
Sources 9 notes
SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.
Models trained on systematically irrelevant traces maintain solution accuracy and sometimes improve out-of-distribution generalization, suggesting traces function as computational scaffolding rather than meaningful reasoning steps.
R1's intermediate tokens carry no special execution semantics and are generated identically to other LLM output. Invalid traces frequently produce correct answers, proving traces are not causally necessary—they correlate with answers via learned formatting, not functional reasoning.
Post-conclusion reasoning—where the model keeps exploring after sufficient evidence for the answer—degrades supervised fine-tuning despite preserving correctness. Removing only this tail improves learning more than removing equally-long random suffixes, proving the harm comes from unnecessary exploration, not length.
Local step-level confidence catches reasoning breakdowns that global averaging masks and enables early stopping before traces complete. This approach achieves comparable accuracy gains to naive majority voting with far fewer generated traces, proving trace quality matters more than quantity.
Show all 9 sources
LEAP demonstrates that models achieve better performance on reasoning and math tasks by intentionally erring on few-shot examples, reflecting on mistakes, and deriving explicit task-specific principles—without additional labeled data or fine-tuning.
Anchoring LLM-judges to reference answers with explicit usage instructions improved judge accuracy by 6.8% and enabled self-improvement training via DPO to match finetuned reward model performance on AlpacaEval and Arena-Hard benchmarks.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Large Language Models Cannot Self-Correct Reasoning Yet
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think