SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can decoding-time tuning preserve knowledge better than weight fine-tuning?

Explores whether applying alignment signals at inference time rather than modifying model weights can better preserve the factual knowledge learned during pretraining while still achieving alignment goals.

Synthesis note · 2026-02-22 · sourced from Training Fine Tuning

Proxy-tuning fine-tunes a small model, then applies the difference between the small tuned and small untuned model's predictions to shift a large untuned model's outputs at decoding time. The large model's parameters are never modified. The method closes 91% of the performance gap between Llama-2-13B and its directly tuned CHAT version, and 88% for the 70B model.

The critical finding: on knowledge-intensive tasks, proxy-tuning sometimes surpasses the performance of direct instruction-tuning. This is because direct fine-tuning modifies model weights — and some of those modifications overwrite pretrained knowledge. Since Why does reasoning training help math but hurt medical tasks?, weight modification risks corrupting the knowledge storage that proxy-tuning leaves intact.

Proxy-tuning primarily promotes reasoning and stylistic tokens. Analysis of the token-level distributional shift shows the largest influence on tokens associated with reasoning patterns and output style — consistent with evidence that "alignment mainly affects style rather than knowledge." This aligns with Does instruction tuning teach task understanding or output format? and Can imitating ChatGPT fool evaluators into thinking models improved?: what fine-tuning actually changes is output distribution, not capability. Proxy-tuning achieves this distributional change without touching the model weights that encode knowledge.

For domain adaptation, proxy-tuning Llama-2-13B using CodeLlama-7B produces 17-32% improvement on coding benchmarks. The small expert provides the distributional guidance; the large base model provides the knowledge. An optional hyperparameter controls the amount of guidance, enabling runtime trade-offs between different generation attributes.

This constitutes a fifth paradigm in the How do knowledge injection methods trade off flexibility and cost?: decoding-time adaptation. Zero training cost on the target model, full knowledge preservation, but requires access to base model logits at inference time.

ARGS (Alignment as Reward-Guided Search) provides a complementary inference-time method. Instead of applying a distributional shift from a tuned proxy, ARGS adjusts model predictions at each decoding step using a reward signal directly. Two components: reward-guided scoring (assigns scores to possible continuations) and token selection (selects a continuation based on scored candidates). A tunable weight controls the trade-off between semantic relevance and alignment criteria — setting it to zero recovers standard maximum-likelihood decoding. ARGS enables rapid personalized alignment without retraining: different users can have different reward functions applied at inference time. Together, proxy-tuning (distributional shift from expert delta) and ARGS (reward-guided decoding) suggest a design space where multiple axes of adaptation — domain knowledge, user preferences, task constraints — can each be applied at decoding time through complementary mechanisms. See Can user preferences be learned from just ten questions? for how per-user reward functions can be efficiently constructed.

Inquiring lines that read this note 161

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does preference optimization undermine conversational grounding in language models? How do models learn from self-generated outputs without cascading failures? Why do language models struggle to implement user intent accurately from prompts? Can base models hide emergent misalignment through alignment training? What prediction granularity best trains models to generate reliable reasoning? Do accumulated memories help or hurt continual learning in models? Which reinforcement learning modifications most improve dialogue quality in language models? How does fine-tuning trade off accuracy against reasoning quality? How do curriculum design and feedback approaches affect model learning? Why do training associations persist despite contradictory contextual information? How do neural networks learn compositional structure from training? How do training data quality and composition affect downstream model performance? How does model capacity affect learning performance on diverse downstream tasks? How does diversity prevent model convergence on superficial patterns? Can inference-time computation adaptively substitute for static model capacity? Can AI systems evade safety evaluations through reasoning manipulation? Why do planning and grounding require opposing optimization strategies? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? How do reward models systematically fail to represent diverse human preferences? Can iterative DPO substitute for online RL in studying misalignment? What makes reasoning traces effective supervision even when they're incorrect? Does pretraining establish the ceiling for what reward learning can improve? What limits language model accuracy in evaluating ideas? What are the fundamental limits of prompting for language models? How does RLHF training shape models to prioritize agreement over accuracy? How can persistent memory architectures preserve information across ultra-long contexts? When should retrieval systems decide to fetch new information? Does intelligent routing among smaller models outperform training larger models? What explains the gap between benchmark scores and true reasoning capability? What capabilities differentiate diffusion from autoregressive language models? Can confidence signals reliably detect flawed reasoning in language models? How do reward signal properties affect model reasoning and safety? What prevents LLMs from applying their reasoning knowledge to improve outputs? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can mechanistic interpretability methods reliably reveal what models actually know? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Can AI agents improve their skills through accumulated experience and reuse? Can smaller specialized models match frontier models on key metrics? Can minimal training unlock latent reasoning already present in base models? Can code harness improvements rival direct model scaling for capability?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 192 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

proxy tuning at decoding time preserves pretrained knowledge better than direct fine-tuning by applying the tuning signal as a distributional shift