SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Does richer teacher context hurt student generalization?

When teachers are given more information during distillation, they produce confident but brittle students. Does this trade-off between in-domain wins and out-of-distribution robustness hold across different task distributions?

Synthesis note · 2026-05-18 · sourced from Training Fine Tuning

The self-distillation degradation finding has a clean causal story. When the teacher model is conditioned on richer information — the correct solution, access to a verifier, additional context that humans would not have at inference time — the reasoning trajectories it produces become more confident and more concise. The teacher knows the answer, so it does not bother to express uncertainty mid-trace. The student, distilling toward these traces, inherits the confident style.

The pattern unfolds along two factors: information richness and task coverage. Richer teacher context → confident traces → suppressed epistemic verbalization → faster in-domain optimization. Limited task coverage means the in-domain wins are real and visible; the model gets better at the narrow distribution it was trained on. As task coverage broadens, the missing uncertainty channel becomes a liability — out-of-distribution problems benefit from expressing uncertainty and adjusting accordingly, and the confident-style student no longer has access to that adjustment mechanism.

This produces a counter-intuitive recommendation for distillation pipeline design. Standard intuition: give the teacher as much information as possible so it produces high-quality traces. The finding inverts this: the teacher's traces become too clean, optimized for cases where confidence is warranted, missing the uncertainty markers that help the student handle cases where confidence is not warranted.

A more robust approach lets the teacher operate with less privileged context, producing traces that include the natural pauses and self-corrections of reasoning under uncertainty. The resulting traces are messier, longer, less obviously "polished" — but they preserve the corrective signal that helps OOD performance.

The deeper observation is that style transfer is part of distillation, not just correctness transfer. The student inherits the teacher's reasoning style, including how the teacher handles or hides uncertainty. Teacher conditioning shapes style, and style shapes generalization. Distillation pipelines that optimize teacher conditioning for correctness alone optimize against generalization without realizing it.

Inquiring lines that read this note 55

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What training dynamics and scale trigger emergence of reasoning capabilities? What makes distillation transfer some model capabilities while suppressing others? What training data selection strategies maximize generalization across difficulty levels? How does synthetic data quality and diversity affect downstream model capabilities? Where and how do personality traits reside in language models? How much do training data properties shape model reasoning? What causes reasoning models to fail or wander off track? How does policy entropy collapse constrain scaling of reasoning-focused RL? Can compression size predict model complexity better than parameter count alone? Why does adding new knowledge through fine-tuning degrade existing capabilities? How much does training format versus domain influence reasoning? Does warmth and empathy training systematically degrade model reliability? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Can prompt-based context override biases that were embedded during pretraining? Can memory architectures handle ultra-long context better than attention? Does model confidence reliably signal actual accuracy in practice? Can we reliably detect when models game evaluations? How do pretraining biases affect reward signal effectiveness in RLVR? Why do token-level mechanisms matter for learning to reason?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 112 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

richer teacher context produces more confident shorter student traces — fast in-domain optimization at the cost of OOD robustness