Does instruction tuning teach task understanding or output format?
Exploring whether models trained on instructions actually learn the task semantics or merely learn to match output distributions. This matters because it challenges assumptions about how fine-tuning improves model behavior.
"Do Models Really Learn to Follow Instructions?" creates two devastating controls. First, simplified task definitions that strip all semantic content, leaving only output space information (e.g., "output one of: A, B, C"). Second, delusive examples containing incorrect input-output mappings. Models trained on either achieve comparable performance to models trained on full, correct instructions. A random baseline achieves 42.6% exact-match versus instruction tuning's 43%.
The implication: instruction tuning primarily teaches the model to map its existing capabilities to the expected output format, not to understand or execute the task as described in the instruction. The semantic content of the instruction — what the task is, how to approach it, what constitutes a correct answer — appears largely irrelevant. What matters is the output distribution: how many classes, what format, what vocabulary.
This connects to a broader pattern. Does training data format shape reasoning strategy more than domain? showed a 7.5x stronger effect of format over domain. Can models pass tests while missing the actual grammar? showed that correct outputs can mask reliance on surface heuristics. The instruction tuning finding adds: even explicit instructions about the task are largely ignored in favor of format signals.
A complementary theory from "Are Emergent Abilities just ICL?" (2309.01809) provides the mechanistic explanation: instruction tuning enables "implicit in-context learning" — mapping instructions to the form required for ICL rather than creating new functional abilities. The evidence: purported emergent abilities are explained by a combination of in-context learning, model memory, and linguistic knowledge. The model's sensitivity to minor prompt variations and tendency to hallucinate are inconsistent with genuine emergent functional abilities but consistent with a model that maps prompts to ICL patterns. This reframes safety concerns: if prompts function as "training mechanisms" rather than interfaces to inherent abilities, the safety landscape changes — the risk is in what ICL patterns exist, not in what abilities have "emerged."
The IT Survey (same source) documents the concern from the other direction: "there has been an intense criticism that IT only captures surface-level patterns and styles rather than comprehending and learning the task." Combined with the False Promise finding that model imitation captures style not factuality, a clear pattern emerges: fine-tuning-based adaptation — whether through imitation, instruction tuning, or domain SFT — preferentially captures distributional and formatting information while leaving underlying capabilities largely unchanged. The capability bottleneck is in the base model, not the adaptation method.
Webson & Pavlick (2021) provide the prompting-level parallel. Evaluating 30+ manually written templates and 13 sets of target words across 390+ prompts, they find models learn identically fast from irrelevant or misleading templates as from instructive ones. Models are "much more sensitive to the choice of LM target words as opposed to the meaning of the instruction templates." Instruction-tuned models can be "too robust" — less sensitive to prompt semantics than non-IT equivalents, suggesting IT trains a form of prompt-blindness. This holds from 235M to 175B parameters. The convergence is striking: both the fine-tuning and the prompting literature arrive at the same conclusion from opposite directions — the semantic content of instructions is largely inert, and what transfers is format and output space information.
Inquiring lines that read this note 178
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance help or harm professional skill development?- Why do workers who understand AI generations learn more than those who only use output?
- Why does AI-improved task performance fail to transfer to independent work?
- Does AI-assisted performance transfer to independent task completion?
- Can explicit reflection during AI-assisted work improve transfer of learning?
- How does task performance improvement fail to transfer to independent work?
- Does AI training preserve learning that transfers to independent subsequent tasks?
- Can we measure perceived skill change against actual independent task performance?
- Does extended exoskeleton use eventually produce meaningful skill transfer?
- Can instruction tuning succeed without explicit task understanding?
- How do training objectives shape what a world model actually learns?
- Does partial trace guidance work better than curriculum learning for hard problems?
- How does behavioral fine-tuning differ from factual knowledge encoding in models?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- Why does critique training produce deeper understanding than imitation training?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- Why does imitation learning alone plateau without outcome-based refinement?
- How do complete multi-turn trajectories differ from isolated task examples?
- Can we predict out-of-distribution generalization without access to downstream tasks?
- Can trained models encode programs more complex than their data-generating process?
- What emergent behaviors do models develop when trained on underspecified pedagogical tasks?
- How does action-level decomposition differ from token-level imitation in supervision?
- What training regimes confound surface mechanisms with their actual causes?
- What makes a good in-context learning example for a given task?
- How do finetuning and pretraining improvements differ in their effects on model capabilities?
- Can we reverse the instruction-following deficit through targeted training?
- What distinguishes surface mechanisms from the training regimes that produce them?
- Why does recontextualizing a behavior during training change whether models learn it?
- What specific tasks should evaluate whether models understand pedagogical sequencing?
- How does demonstration coverage in context examples determine operation generalization?
- Why does instruction-tuning reduce a model's context-following behavior?
- Why does fine-tuning function as character training rather than capability training?
- Can granular sub-task training for function calling improve both open and proprietary models?
- How much of the combinatorial task space must training data cover?
- How do task difficulty and skill type interact in model performance?
- Can a single model trained on two tasks predict untrained decision tasks?
- Does fine-tuning actually change model capabilities or only output distribution?
- How much task-similar finetuning data does test-time training actually need?
- Which finetuning method works best across different task and data regimes?
- How do task frequency and complexity interact with model capacity during training?
- Can intentional data-mixture design replace model scaling for rare task learning?
- Can benchmarks designed for shortcut learning detect heuristic override failures?
- How does task contamination differ from test set data leakage?
- What is the gap between benchmark performance and real workplace task completion?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- What distinguishes genuine task improvement from evaluator exploitation?
- Are cheap testbeds and skewed task distributions linked by design necessity?
- Does measured performance gain reflect true task improvement or evaluator exploitation?
- Does semantic auditing of instruction data improve performance uniformly across different model sizes?
- What distortions do automated benchmarks introduce compared to real tasks?
- What drives the mismatch between general benchmark leadership and task-specific performance?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- How does post-training shift models from passive prediction to on-policy action?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- What role does pretraining play in distinguishing system capability from deployed behavior?
- How does the knowing-doing gap widen as tasks become more complex?
- Can models maintain multiple task interpretations simultaneously before committing to a single policy?
- How does stage-wise training scheduling resolve conflicts between constraint-following and creative tasks?
- How do task-agnostic and task-oriented skills differ in coverage and reuse?
- Can curated demonstrations compensate for smaller or simpler training environments?
- Why does mixed instruction data sometimes hurt specific model capabilities?
- What distinguishes instance seeds from full input-output exemplar requirements?
- How much performance is lost when converting pretrained checkpoints versus training from scratch?
- Do identical task structures mean repeated instances or new synthetic samples with same design?
- Can training data organization by capability outperform source or task-based mixing?
- Does alignment training create bidirectional instruction and response mappings?
- Does correct model behavior guarantee internal alignment of learned objectives?
- What specific behavioral patterns should alignment examples target for maximum effect?
- Does pretraining poisoning at scale persist through instruction alignment?
- Can mechanistic interpretability tools decode the biases alignment training conceals?
- What makes principle-response mutual information sufficient for behavioral alignment?
- Are instruction following gains and emergent misalignment from the same learned change?
- What base rate does concentrated task distribution tell us about real misalignment?
- Do verbal alignment benchmarks measure representation or just output compliance?
- Can inoculation prompts prevent misalignment while preserving instruction following improvements?
- Can prompting unlock compositional skills that pretraining already learned?
- Can dynamic instance-specific prompt selection solve the generalization problem across tasks?
- Can demo placement be tuned as a task-specific hyperparameter?
- Why do primacy effects peak at specific instruction densities?
- Does input length alone explain instruction density performance loss?
- Are instruction-tuned models more or less sensitive to prompt semantics than others?
- How does explicit exploratory prompting compare to fine-tuned reinforcement learning for in-context adaptation?
- How do prompting and activation steering relate as compression strategies?
- How do input-side defenses separate task methodological and framing intents?
- How much does instruction prompt design control what alignment target an AI annotator enforces?
- Can persistent prompt optimization encode a scoring shortcut into reused instructions?
- Can steering internal features bypass or override prompt-level instructions in simulations?
- What execution feedback signals drive context updates without supervision labels?
- Can a separate mediator layer improve intent understanding before task execution?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Why do instruction following and reasoning capability trade off in training?
- Can reasoning fine-tuning improve both capability and instruction compliance together?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- How do procedural versus factual knowledge differ in pretraining versus fine-tuning?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- Can fine-tuning ever teach semantic inference instead of amplifying training shortcuts?
- Does fine-tuning models for specific tasks destroy their ability to reason?
- Do instruction-tuned models learn tasks or just output format distributions?
- Why does instruction tuning hurt knowledge-intensive tasks more than reasoning tasks?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Does training on granular tasks beat training on the full function calling problem?
- Can extracted skills transfer effectively across different domains and model architectures?
- Where does skill extraction fail compared to genuine model adaptation?
- Can training on diverse related tasks be more efficient than task-specific training?
- Do text-space skills transfer learning across different frontier models?
- Can in-context learning replicate the timing effects that RL teaches models?
- Does format-based pretraining determine how models respond to reinforcement learning?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- How does preference-based training compare to supervised fine-tuning for function calling?
- How does task-oriented fine-tuning compare to preference tuning methods?
- Does approaching human performance mean learning the same grammatical rules?
- Do instruction-tuned models prefer conversational over formal source language?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Does scaling reasoning capability create tradeoffs with instruction following?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Why does target probability matter more than task logical complexity?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Why does stronger reasoning reduce model compliance with instructions?
- Why does instruction-following capability decrease as models scale stronger?
- Why do more capable reasoning models become harder to control by instruction?
- Does highlighting input features reduce human over-reliance on machine outputs?
- Does foundational model training or user priors more strongly shape final outputs?
- Do negative constraints require fundamentally different training signals than positive instructions?
- Why do vector embeddings fail to measure task relevance in production RAG?
- Can vector embeddings measure task relevance instead of semantic similarity?
- How do instruction backtranslation and MAGPIE demonstrate self-generation principles?
- Do self-generated explanations outperform passive instruction for oversight?
- What makes high-quality GUI instruction data different from general vision data?
- Why does identifying UI element types and locations enable downstream task learning?
- What distinguishes task-specific heuristics from genuine world models?
- What is the difference between changing model outputs versus changing internal representations?
- How does belief-behavior inconsistency relate to instruction execution splits?
- What does leveraging internal representations during training actually mean operationally?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- What distinguishes data that generalizes broadly from task-specific memorization?
- Why does specializing to one task make future task learning harder?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- Can we predict which tasks will decompose into modular subnetworks?
- Can curvature measurements predict task difficulty without behavioral labels?
- What limits the extrapolation of learned operations like rotation and reflection?
- Do legitimate task signals exploit the same position and framing vulnerabilities as attacks?
- Can verifiable execution traces replace fluent output as a training signal?
- Do weight-space skills lose detail compared to textual skill descriptions?
- Why do generic skill descriptions evolve into execution-oriented ones?
- Why do imitation learning agents stay locked within human demonstration patterns?
- Why do benchmark tasks differ from real occupational workflows in practice?
- Why do interactive tasks show larger capability gaps than bounded tasks?
- Does synthetic fine-tuning create evaluation awareness similar to natural model reasoning?
- What distinguishes a general evaluation direction from task-specific behavioral patterns?
- How does supervised finetuning amplify evaluation awareness in base models?
- How does instruction tuning affect evaluation detection more than model scale?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does training data format shape reasoning strategy more than domain?
What explains why models trained on multiple-choice data reason differently than those trained on free-form text? The research isolates format and domain effects to measure which one matters more.
format > domain at 7.5x; this adds format > instruction semantics
-
Can models pass tests while missing the actual grammar?
Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.
same mechanism in linguistic domain
-
Can small models reason well by just learning output format?
Does reasoning performance depend primarily on adapting how models express outputs rather than acquiring new knowledge? The Tina research tests this by applying LoRA to a 1.5B model during reasoning training.
LoRA as format adapter aligns with IT as format teacher
-
Does supervised fine-tuning actually improve reasoning quality?
While SFT boosts final-answer accuracy, does it degrade the quality and informativeness of the reasoning steps that justify those answers? This matters for high-stakes domains requiring auditable decision-making.
SFT raises accuracy because it teaches the output format, not because it improves reasoning
-
Why do chain-of-thought examples fail across different conditions?
Chain-of-thought exemplars show surprising sensitivity to order, complexity level, diversity, and annotator style. Understanding these brittleness dimensions could reveal what makes reasoning prompts robust or fragile.
complementary evidence of format-over-substance: IT achieves accuracy through format matching alone, while CoT exemplar brittleness shows reasoning performance depends on surface exemplar properties (order, style, complexity) rather than semantic content
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning
- A Survey on Post-training of Large Language Models
- Exploring Format Consistency for Instruction Tuning
- It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- LESS: Selecting Influential Data for Targeted Instruction Tuning
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- Instruction Induction: From Few Examples to Natural Language Task Descriptions
Original note title
instruction tuning teaches output format distribution not task understanding — simplified and delusive instructions achieve comparable performance