If an AI learns from everyone's preferences at once, does better data alone stop it from quietly favoring the majority?
Does reward model training data quality determine fair preference aggregation outcomes?
This explores whether cleaner or better training data for reward models (the systems that score AI outputs during RLHF) would be enough to fairly balance the preferences of different people, or whether the unfairness comes from somewhere else.
This explores whether the fairness problem in reward models is really a data-quality problem, that is, whether better labels would let one reward model balance many people's preferences fairly. The corpus mostly says no. Even perfect data can't fix the main failure, because it comes from the model's shape, not the data. Can a single reward model represent diverse human preferences? proves that fitting a single reward model to pooled preferences quietly erases minority viewpoints, and that this follows mathematically. Can aggregate reward models satisfy genuinely disagreeing users? gives a concrete case. If users split 51–49, a single score has to either leave 49% unhappy every time or leave everyone unhappy half the time. That note calls it "a representational failure, not a quality problem." A single number has no way to hold a real disagreement.
Data quality still matters, though, in a less obvious way. Do all annotation responses measure the same underlying thing? draws on behavioral science to show that annotation responses aren't all the same kind of thing. Some are genuine preferences. Some are "non-attitudes," where the annotator didn't really care and picked something. Some are preferences people made up on the spot because the question asked for one. If you treat all three as real preferences, the aggregate gets distorted before any fairness math begins. So the problem has two layers: you need the right structure to represent disagreement, and you need to know which disagreements are real in the first place.
One proposed fix is to stop forcing everyone into one score. MaxMin-RLHF learns a mixture of preference groups and optimizes for the worst-off group, an idea borrowed from social choice theory. Other work personalizes instead. Can user preferences be learned from just ten questions? infers a user's personal weighting from about ten well-chosen questions. Can text summaries beat embeddings for personalized reward models? uses readable text summaries of what a user prefers in place of opaque embedding vectors. Personalization has its own cost, though. Does personalizing reward models amplify user echo chambers? warns that removing the averaging effect also removes a brake: a reward model tuned to one person can learn to flatter them and reinforce polarization, much as recommender systems have done.
A result from outside RLHF adds something easy to miss: training data isn't a neutral snapshot of preferences. Why do ranking systems need to model selection bias explicitly? shows that YouTube's ranking data partly reflects what the system already chose to show. Without explicit correction, the model amplifies its own past decisions. Preference data collected from a deployed AI has the same risk, since people can only rate outputs the model already produces. Does reward hacking always stem from the same failure? puts it generally: optimizing hard against any scoring signal that only partly captures what you want will exploit the gap.
So the answer is that data quality affects fairness but doesn't decide it. The bigger factor is a design choice: one score or many, average or protect the worst-off group, and who counts as the population. Better data makes each of those choices work better, but it can't stand in for making them.
Sources 8 notes
MaxMin-RLHF proves an impossibility result: fitting one reward model to aggregated preferences silently erases minority viewpoints. The solution is learning a mixture of preference distributions and optimizing a MaxMin objective from social choice theory to protect the worst-off groups.
Single reward models trained on aggregated preferences cannot represent disagreement. A 51-49 preference split forces a choice between leaving 49% unhappy always or leaving everyone unhappy half the time. This is a representational failure, not a quality problem.
Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.
PReF learns base reward functions from preference data, then uses active learning to select maximally informative questions that reduce coefficient uncertainty. Users can be personalized via inference-time reward alignment without weight modification.
PLUS trains summarizers and reward models jointly, learning that text-based preference summaries capture dimensions zero-shot summaries miss. These summaries transfer to GPT-4 for zero-shot personalization and remain interpretable to users.
Show all 8 sources
Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.
YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.
Reward hacking arises during weight training, output selection, and prompt revision through a shared failure: optimization against signals that incompletely represent the actual task. The substrate matters less than the misalignment between the scoring function and ground truth.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Capturing Individual Human Preferences with Reward Features
- Measuring Human Preferences in RLHF is a Social Science Problem
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
- Enhancing personalized multi-turn dialogue with curiosity reward
- Beyond Preferences in AI Alignment
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Language Model Personalization via Reward Factorization
- Personalized Language Modeling from Personalized Human Feedback