SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Why do language models correct user errors but not their own?

When models see identical errors attributed to users versus themselves, they fix the former but not the latter. Is this a knowledge gap or a learned blind spot that could be fixed?

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

Self-Correction Bench isolates whether a model's failure to fix an error is a knowledge problem or an activation problem, by injecting "the same error as either an external (user-attributed) or internal (model-attributed) error, keeping all other context identical." Testing 14 open-source non-reasoning models, it reports "a 64.5% Self-Correction Blind Spot: models correct external errors but fail on identical internal ones, proving the capability exists but is not activated." The gap holds "across model families, scales, and task complexities ranging from trivial arithmetic to multi-step mathematical reasoning," and is validated on closed-source frontier models and non-mathematical domains (logic, object tracking). On models' own naturally generated errors, rather than injected ones, "at least 4.3–10.8% of the errors a model commits are ones it demonstrably had the knowledge to catch" — a lower bound, not the full blind spot, since on-policy errors conflate knowledge and activation.

The paper traces the cause to post-training data composition: "supervised fine-tuning datasets lack error-correction sequences" because human demonstrations and preference data "strongly favor polished, error-free responses," and synthetic instruction data inherits the same bias from the human data it was distilled from. Fine-tuning on "as few as 5,306" error-correction traces, two epochs of LoRA, reduces the blind spot by 76.0%, showing the gap is a trainable artifact rather than a hard ceiling. Mechanistically, the authors isolate "a transferable conversational-role direction in representation space that causally gates self-correction," demonstrated in two model families (Llama and Qwen) at 7–8B scale. Separately, appending "Wait" — a training-free, test-time intervention — cuts the blind spot by 89.3% through "a nearly independent pathway," implying at least two separable triggers for correction: a training-induced role attribution, and an inference-time metacognitive one.

Splitting "capability exists but isn't used" from "capability doesn't exist" echoes Why do language models lose performance in longer conversations?, which locates the Lost-in-Conversation drop in an alignment gap rather than a model limitation; both treat a measured deficit as a trained-in disposition, not a capability ceiling. This also gives a second cause for SFT's documented failure to teach self-correction: Why does self-correction training on offline data fail? attributes that failure to training on the wrong error distribution, while Self-Correction Bench attributes it to near-total absence of correction sequences in the data at all — a data-composition gap its own LoRA result suggests can close without RL. It also names a mechanism distinct from Do models fail worse when their own errors fill the context?: that note locates internal-error persistence in error accumulation over a long horizon, while this paper locates it in a single role-attribution switch present even in a one-shot setting with a single injected error.

The paper is explicit that "the magnitude is a property of the controlled design rather than an estimate of how often models fail to correct themselves in deployment" — 64.5% describes an injected-error benchmark, not a deployment rate; the only deployment-relevant figure given is the smaller 4.3–10.8% lower bound on naturally occurring errors. The conversational-role direction is demonstrated in exactly two model families at 7–8B scale, with universality beyond them untested, and all 14 models are non-reasoning models — whether reasoning models, whose training data already contains more correction sequences per the paper, show the same blind spot at the same magnitude is not established here.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do models learn from self-generated outputs without cascading failures? How do users confuse explanation quality with actual system accuracy?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 125 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms correct the same error when it is attributed to the user but not when it is attributed to themselves — a 64.5 percent self-correction blind spot