SYNTHESIS NOTE
Topics›RAG›this note

Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?

Explores whether rewarding coherent reasoning patterns during training helps models internalize domain knowledge better than standard fine-tuning approaches that treat all tokens equally.

Synthesis note · 2026-02-22 · sourced from RAG
RAG

SFT on domain knowledge treats all tokens equally. A training example of a medical question answered correctly does not distinguish between the tokens that encode critical clinical reasoning and the tokens that are boilerplate formatting. CPT (continual pre-training) is worse: it processes entire domain documents without targeting clinically critical information. Both approaches fail at knowledge coherence — the model may learn isolated facts without integrating them into the connected knowledge structures needed for complex reasoning.

RLAG (Reinforcement Learning from Augmented Generation) takes a different approach. For each question, generate two responses: one with retrieved domain context as prefix, one without. The augmented response is the "preferred" response (the model sees what the correct answer looks like with evidence support). The unaugmented response is what the model can produce from parametric knowledge alone. The reward signals: answer accuracy and explanation rationality — not just whether the final answer is right but whether the reasoning that produced it is coherent.

The iterative cycle: sample → compute rewards → optimize → repeat. With each cycle the model internalizes the knowledge patterns from retrieved context, gradually reducing the gap between its unaugmented performance and augmented performance. The retrieved context during training becomes scaffolding that the model eventually internalizes.

The key difference from SFT: RLAG rewards the model for the quality of its knowledge representations, not just for reproducing training examples. A model that gets the right answer through incoherent reasoning is not rewarded. A model that produces a coherent explanation from genuinely integrated knowledge is.

This adds a new mechanism to the How do knowledge injection methods trade off flexibility and cost?: RL-from-augmentation is not purely dynamic (inference-time RAG) nor purely static (SFT/CPT) — it uses dynamic context during training to progressively embed what it learned into weights, creating models that can reason coherently without retrieval at test time.

Inquiring lines that read this note 84

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI agents improve their skills through accumulated experience and reuse? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? How does fine-tuning trade off accuracy against reasoning quality? Which reinforcement learning modifications most improve dialogue quality in language models? How do real-world evaluations reveal AI capabilities that benchmarks hide? What are the fundamental limits of prompting for language models? Can smaller specialized models match frontier models on key metrics? What enables conversational agents to guide rather than just respond? Does pretraining establish the ceiling for what reward learning can improve? Can inference-time computation adaptively substitute for static model capacity? Does preference optimization undermine conversational grounding in language models? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Do accumulated memories help or hurt continual learning in models? How do curriculum design and feedback approaches affect model learning? Can minimal training unlock latent reasoning already present in base models? What prediction granularity best trains models to generate reliable reasoning? Can confidence signals reliably detect flawed reasoning in language models? How do reward signal properties affect model reasoning and safety? How do reward models systematically fail to represent diverse human preferences? What makes process supervision effective for training complex reasoning models? How do agents learn to distinguish valuable feedback from noise? Can artificial systems establish authority in domains requiring expert judgment? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 133 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl from augmented generation embeds domain knowledge more effectively than sft by rewarding coherent knowledge structures