Can online AI feedback make preference alignment truly on-policy?
Direct preference optimization methods like DPO suffer from off-policy misalignment because preference data comes from older models. Can sampling feedback from the current model each iteration solve this drift and improve alignment quality?
Direct alignment from preferences (DPO, IPO, SLiC) is attractive because it skips the separate reward model and updates the policy directly from pairwise preferences. But its preference datasets are collected ahead of training and never updated, and the responses usually come from a different model — so as the policy evolves, alignment becomes inevitably off-policy and prone to overfitting. OAIF's fix is simple: on each training iteration, sample two responses from the current model and prompt an LLM annotator to pick the preferred one, supplying online feedback. Despite its simplicity, human evaluation shows OAIF beats both offline DAP and RLHF, and it mitigates reward over-optimization — the overfitting that plagues offline DAP.
Two keepers. First, the online vs offline distinction matters more than the choice among DAP variants: OAIF improves DPO, IPO, and SLiC alike, isolating on-policy feedback as the lever. Second, the AI annotator's feedback is controllable via instruction prompts — you can steer the alignment target by changing how you ask the judge to choose.
This connects the vault's alignment-method thread to the LLM-as-judge thread. The controllable AI annotator inherits the risks documented in Can LLM judges be fooled by fake credentials and formatting? — an online judge that is biased steers the policy toward those biases — and the on-policy framing rhymes with Can agents learn from failure without updating their weights? in treating fresh, current-model feedback as the signal that matters.
Inquiring lines that read this note 13
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does alignment training create genuine alignment or just output compliance? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Can iterative DPO replicate online reinforcement learning dynamics for research?- What misaligned behaviors does iterative DPO produce compared to online RL?
- Does environment choice explain differences between iterative DPO and RL results?
- Does iterative DPO generalize like online reinforcement learning?
- What specific properties of online RL does iterative DPO actually preserve?
- How many rounds of iterative DPO are needed to induce misalignment behaviors?
- Does iterative DPO generalize identically to online RL on the same benchmark tasks?
- Does iterative DPO preserve the same misalignment mechanisms as online reinforcement learning?
- Does iterative online DPO fidelity match true reinforcement learning for safety research?
- How does iterative DPO compare to full RLVR for studying emergent misalignment?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
OAIF's controllable AI annotator inherits LLM-judge biases that then steer the policy
-
Do unimodal reward models actually serve all user preferences?
Standard RLHF assumes a single utility function across all users, but what happens when preferences genuinely conflict? Does averaging these opposing preferences into one model systematically fail certain groups?
adjacent alignment-method concern: OAIF is on-policy single-judge; diverse-preference work questions the single-judge assumption
-
Can iterative DPO replace reinforcement learning for studying reward hacking?
Researchers propose using iterative direct preference optimization instead of costly reinforcement learning to study how models develop misalignment through reward hacking. The question is whether this cheaper method preserves the important properties that make the original approach scientifically valuable.
a safety-research use of the same online-versus-offline DPO distinction: a semi-online loop proposed as a cheap stand-in for RL, where the on-policy reading of "iterative" is the vault's inference and fidelity to online RL is not compared
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Direct Language Model Alignment from Online AI Feedback
- Improving Small-Scale Large Language Models Function Calling for Reasoning Tasks
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Self-Improving Model Steering
- Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
- Bridging Offline and Online Reinforcement Learning for LLMs
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- KTO: Model Alignment as Prospect Theoretic Optimization
Original note title
online AI feedback makes direct preference optimization on-policy — sampling from the current model and judging with an LLM beats offline DPO and RLHF