INQUIRING LINE

When someone keeps giving the same answer about what they want, do they really want it — or did the question just make them say so?

Does measured opinion stability reflect genuine preference or elicitation artifact?

This explores whether a person giving the same answer about their preferences again and again means they really hold that view, or whether the steadiness is produced by how the question is asked, and why that matters for training and judging AI.


This explores whether stable, repeatable answers about what people prefer show a real underlying preference, or whether the way we ask creates the steadiness. The corpus has a pointed answer: it can be either, and the only way to tell is to change how you ask and see what holds up. Behavioral science, drawing on about sixty years of survey research, sorts responses into three kinds. Genuine preferences stay put when the question changes. Non-attitudes are answers people give because they were asked, not because they held a view. Constructed preferences are made up on the spot from whatever the question's framing supplies Do all annotation responses measure the same underlying thing?. Stability within a single setup tells you very little. Stability across different setups is the real signal.

This matters because RLHF, the main way chatbots are tuned to human taste, mostly skips the check. It treats every annotator click as real signal, so reward models end up learning survey artifacts as if they were human values Are RLHF annotations actually measuring genuine human preferences?. The argument is that measurement comes before aggregation. Before asking how to combine many people's preferences fairly, you have to ask whether you measured a preference at all.

The same trap shows up in an unexpected place: the models themselves. Setting an LLM's temperature to zero makes it give the same answer every time, but that answer is still one draw from its range of possible outputs. Repeatability is not reliability Does setting temperature to zero actually make LLM outputs reliable?. People and models fail the same way here: a fixed procedure can make noise look consistent.

Answers can also be steady because of outside forces rather than inner conviction. Online ratings are pulled toward earlier ratings, and that pull compounds over time Do online ratings actually reflect independent customer opinions?. Which recommender put two products side by side changes who sees them, which in turn changes whether their ratings converge Do different recommender types shape opinion convergence differently?. In debates, what voters already believed predicts outcomes better than anything the debaters said Does what readers believe matter more than what debaters say?. Some 'preferences' are not about the content at all. In Search Arena, users rewarded answers with more citations nearly as much when the citations were irrelevant Do users trust citations more when there are simply more of them?.

Aggregation is not useless, though. Chatbot Arena's crowd votes match expert rankings, partly because the questions are varied and discriminating Can crowdsourced votes reliably rank language models?. Averaging across many people and prompts can wash out individual artifacts. That is also why personalized reward models are risky: once you fit to one person, nothing averages away their non-attitudes, and the system can learn to flatter them Does personalizing reward models amplify user echo chambers?. The surprise here is that the push toward AI that reflects your own preferences most closely may be the one most likely to encode the parts of you that were never real preferences.


Sources 9 notes

Do all annotation responses measure the same underlying thing?

Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.

Are RLHF annotations actually measuring genuine human preferences?

Sixty years of behavioral science evidence shows humans produce survey responses without genuine underlying preferences. RLHF ignores this, training reward models on non-attitudes and constructed preferences as if they were stable signal.

Does setting temperature to zero actually make LLM outputs reliable?

Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.

Do online ratings actually reflect independent customer opinions?

Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.

Do different recommender types shape opinion convergence differently?

Research shows that frequently-bought-together and co-viewed recommendation networks produce different opinion convergence patterns. The mechanism: each recommender type attracts different audience segments with different prior expectations, shaping both who sees products together and how they rate them.

Show all 9 sources
Does what readers believe matter more than what debaters say?

Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Does personalizing reward models amplify user echo chambers?

Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.