Can a reasoning model fix its own mistakes just because you tell it 'wait, keep thinking'?
Does appending a single word at test time unlock model self-correction abilities?
This explores whether a tiny test-time nudge, such as adding a word like "Wait" to make a model keep thinking, can make a model catch and fix its own mistakes, or whether self-correction needs something deeper.
This explores whether a tiny test-time nudge, such as appending "Wait" so a model keeps reasoning instead of stopping, can make a model genuinely fix its own errors. The collection has no paper that tests that exact trick head-on. It does have a lot of evidence on the question underneath it: when a model is prompted to reconsider, does the reconsidering actually correct anything? That evidence leans clearly toward no.
Start with what happens when reasoning models reflect on their own. Across eight reasoning models, reflections rarely changed the answer. They mostly confirmed what the model had already decided, so the "second look" works more as a ritual than as a repair Is reflection in reasoning models actually fixing mistakes?. Forcing extra revision can make things worse. In QwQ, R1 and LIMO, most revisions kept wrong answers wrong, smaller models often switched correct answers to incorrect ones, and longer chains with more revisions went with lower accuracy Does self-revision actually improve reasoning in language models?. If a single word buys more thinking time, these results suggest it mostly buys more confirmation, and sometimes it gives the model more chances to talk itself out of a right answer.
The deeper reason is that a model judging its own work isn't a neutral checker. Models over-trust answers they generated themselves, because a high-probability answer also *feels* correct when the same model evaluates it Why do models trust their own generated answers?. That fits the broader finding that without an outside signal telling the model it got something wrong, self-correction lowers reasoning accuracy across GPT-3.5, GPT-4 and Llama-2. Even multi-agent debate turned out to be majority voting under another name Can language models fix their own reasoning mistakes?. A prompt word can't supply the missing ingredient, which is information the model doesn't already have.
So where does real self-correction come from? The corpus points to training, not prompting. Models learn to correct mistakes when they practice on their *own* errors through online RL. Fine-tuning on pre-made correction examples fails because those mistakes don't look like the ones the model actually makes at test time Why does self-correction training on offline data fail?. Scale matters too: a 1-trillion-parameter model discovered self-verification on its own under plain RL, while a 104B model needed hand-designed rewards to get there Does scale alone teach models to reason without hand-crafted rewards?. That suggests a nudge can only draw out a checking habit the model already has.
That last point is what makes the "one word" idea interesting rather than simply wrong. A cheap trigger can work as a key, but only if training has already put the habit behind the door. Related work shows something similar during training: about 1,000 examples of "how to deepen shallow reasoning" can activate latent ability that then supports repeated self-improvement Can models improve themselves on tasks without verifiable answers?. The better question is less "does the word work?" and more "what did the model learn before the word arrived?"
Sources 7 notes
Analysis of 8 reasoning models shows reflections rarely change answers and primarily serve as post-hoc confirmation. Training on longer reflection chains improves first-answer quality, not self-correction capability.
Evidence from QwQ, R1, and LIMO shows most revisions retain wrong answers rather than correcting them. Smaller models frequently switch correct answers to incorrect during revision, and longer chains with more revisions correlate with lower accuracy.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Across GPT-3.5, GPT-4, GPT-4-Turbo, and Llama-2, self-correction without external labels degrades reasoning accuracy. Multi-agent debate gains match plain self-consistency at identical cost, suggesting debate is consistency voting, not genuine correction.
SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.
Show all 7 sources
Ring-Zero found a scale threshold where pure zero RL becomes sufficient: a 104B model required hand-designed rewards for structured reasoning and self-verification, but a 1T model discovered these strategies autonomously. This suggests reasoning-scaffolding research has value tied to model size.
Training on just 1000 examples of reasoning enrichment—showing how to expand shallow reasoning into deeper thought—enables models to iteratively improve on general tasks without external verification. The catalyst data activates latent reasoning ability and provides a stable signal across multiple improvement iterations.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Large Language Models Cannot Self-Correct Reasoning Yet
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
- Can Large Reasoning Models Self-Train?
- Training Language Models to Self-Correct via Reinforcement Learning
- Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies