SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can splitting adaptation into two channels reduce forgetting?

When language models adapt to new tasks, does separating task-specific learning (via prompt context) from persistent parameter updates help preserve both generalization ability and the model's original capabilities?

Synthesis note · 2026-05-28 · sourced from Training Fine Tuning

Treating parameter updates as the sole mechanism of adaptation creates a bottleneck: every improvement — a reusable reasoning skill, a task heuristic, even a transient lesson from recent rollouts — has to be written into the same persistent weights. Because the whole policy lives in those weights, any update that raises in-domain reward simultaneously drags the model away from its base behavior, reducing entropy, hurting out-of-distribution generalization, and eroding the model's ability to adapt to future tasks (plasticity loss).

Fast-Slow Training resolves this by refusing to make weights carry everything. It splits adaptation into a slow parametric component (model weights, expensive to update, persisting long-lived behavior) and a fast textual component (prompts, instructions, task context, optimized via reflective prompt evolution with GEPA). The fast channel absorbs task-specific and rapidly-changing information from textual feedback; the slow channel consolidates only persistent behavior and stays closer to the base model. Interleaving the two — RL updates plus context optimization — reaches matched performance with 1.4–3x fewer optimizer steps and a higher asymptote, while leaving the model far closer to its origin.

Why it matters: it reframes catastrophic forgetting as a misallocation problem rather than an inherent cost of learning. Forgetting happens because we force weights to store things that did not belong in weights. Route the transient and task-specific into context, and the weights stay general — so there is less to forget. This is a division-of-labor argument: the two channels operate at different timescales (an echo of System 1 vs System 2) and each does what it is suited for. The counterpoint is that the fast channel's capacity is bounded by context length and prompt-optimization quality, so genuinely large bodies of new knowledge still have to land in weights eventually.

Inquiring lines that read this note 57

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does adding new knowledge through fine-tuning degrade existing capabilities? Can memory architectures handle ultra-long context better than attention? What role does sparsity play in model behavior and scaling decisions? How does AI adoption across firms reshape employment and inequality? What makes distillation transfer some model capabilities while suppressing others? Why can't prompting alone inject genuinely new knowledge into models? How does decomposing tasks improve reasoning and prevent failure propagation? Can prompt-based context override biases that were embedded during pretraining? What capability trade-offs arise from domain specialization through fine-tuning? What causes retrieval-augmented generation systems to fail despite access to external knowledge? What structural properties of attention create systematic model biases? How should agents manage memory granularity to improve long-term performance? How should inference compute be allocated based on problem difficulty? How should systems decide whether to retrieve or reason alone? How much do training data properties shape model reasoning? Why does memory consolidation cause performance regression in continual learning? Is reasoning capability latent in base models or created by post-training? How does synthetic data quality and diversity affect downstream model capabilities? Why do token-level mechanisms matter for learning to reason? Does AI assistance promote real skill development or substitute for independent learning? What mechanisms preserve shared understanding in evolving conversations? How do surface patterns enable correct outputs but reduce robustness?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 155 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

splitting adaptation into slow weights and fast textual context avoids catastrophic forgetting and plasticity loss