SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can editing hidden representations beat weight updates for finetuning?

Does intervening directly on a frozen model's representations offer a better path to parameter-efficient adaptation than current weight-based methods? This challenges the dominant PEFT paradigm by treating representations as the semantic lever instead.

Synthesis note · 2026-06-03 · sourced from Training Fine Tuning

Parameter-efficient finetuning (PEFT) adapts large models by updating a small number of weights (LoRA and variants). ReFT starts from a different premise drawn from interpretability: representations encode rich semantic information, so editing representations might be more powerful than editing weights. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. Its strong instance, LoReFT (low-rank linear subspace ReFT), is a drop-in PEFT replacement that is 10–50× more parameter-efficient than prior state-of-the-art PEFTs and almost always outperforms them across eight commonsense-reasoning, four arithmetic-reasoning, instruction-following (Alpaca-Eval), and GLUE tasks.

The keeper is the conceptual bridge: interpretability findings (that meaning lives in representations as directions/subspaces) become an adaptation method — intervene in the representation subspace rather than perturb weights. This unifies steering and finetuning: the same handle used to interpret a model can be used to adapt it.

This connects the vault's PEFT and mechinterp threads. It operationalizes the linear-representation premise behind Can dictionary learning scale to production language models? (features as steerable directions) as a finetuning technique, and it rhymes with Does reinforcement learning update only a small fraction of parameters?: adaptation concentrates in a low-dimensional subspace, whether of weights or representations.

Inquiring lines that read this note 39

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How much do training data properties shape model reasoning? How do surface patterns enable correct outputs but reduce robustness? Why does adding new knowledge through fine-tuning degrade existing capabilities? What capability trade-offs arise from domain specialization through fine-tuning? What enables genuine semantic understanding in language models? What role does sparsity play in model behavior and scaling decisions? How do neural networks achieve compositional generalization at scale? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How do agent-learned skills transfer and improve across different tasks? How do training data properties determine the emergence of internal misalignment? How does harness optimization generalize across different model architectures and domains? How should inference compute be allocated based on problem difficulty? What makes distillation transfer some model capabilities while suppressing others? Can prompt-based context override biases that were embedded during pretraining? How should retrieval systems handle complex multi-step reasoning? Can reasoning scale in latent space without tokens? How do false presuppositions and sycophancy drive persistent false beliefs in models? Can mechanistic interpretability reliably guide practical model design choices? What fundamental constraints limit how effectively agents can improve themselves? Why can't prompting alone inject genuinely new knowledge into models? Does alignment training create genuine alignment or just output compliance?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 161 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

representation finetuning intervenes on frozen hidden representations instead of weights and is far more parameter-efficient than LoRA