SYNTHESIS NOTE
Topics›Alignment›this note

Do all annotation responses measure the same underlying thing?

Explores whether RLHF's treatment of all annotations as equivalent signals overlooks fundamental differences in what those responses actually represent—stable preferences versus non-attitudes versus context-dependent constructions.

Synthesis note · 2026-04-07 · sourced from Alignment

Behavioral science's six-decade accumulation of preference elicitation research produces a taxonomy that RLHF practice collapses into a single signal. The three categories matter because they require different treatment — and treating them uniformly is the upstream mistake that Are RLHF annotations actually measuring genuine human preferences? argues contaminates the entire pipeline.

Genuine preferences manifest stably across equivalent measurement conditions. Ask the same question with different surface wording, different framing, different order, and the response stays the same. This is what the reward model is supposed to be learning. Only this category is safe to aggregate in the way standard RLHF aggregates.

Non-attitudes are responses generated to satisfy the question without any stable underlying opinion. The respondent has never formed a view on the matter, but the measurement protocol demands an answer, so one gets produced. Non-attitudes are especially pervasive for value-laden questions — precisely the questions that matter most for alignment. Non-attitudes look like genuine preferences in a single measurement but fail the consistency test: re-ask the same respondent and you get a different answer because there was never a stable view to retrieve. Current RLHF treats these as noise to filter or minority views to downweight. The behavioral science view is different: non-attitudes contain no signal at all and should be excluded, not averaged with genuine preferences.

Constructed preferences are assembled on the spot from contextual cues and framing. The respondent is not uncertain (as in a non-attitude); they are producing a coherent answer that depends on the measurement context. Change the context — different anchoring, different comparison class, different framing — and you get a different coherent answer. This category carries real information, but about the interaction between person and context, not about a stable property of the person. RLHF treats constructed preferences as context-independent preferences and trains reward models on them as if they were. The result: reward models that look good on in-distribution evaluation but fail when the deployment context differs from the annotation context.

Measurement artifacts form a fourth related category: same question measuring different constructs for different respondents. One annotator interprets "helpful" as "completes the task"; another interprets it as "gives correct information even when unasked"; a third interprets it as "avoids making the user feel incompetent." They provide coherent, stable responses — each tracking a real preference of theirs — but they are not tracking the same thing. RLHF aggregates them as if they were.

The diagnostic criterion that separates these is consistency across equivalent measurement conditions. Genuine preferences pass; non-attitudes, constructed preferences, and measurement artifacts each fail in distinctive ways. Non-attitudes fail on re-ask (no stable view). Constructed preferences fail on context perturbation (context-dependent). Measurement artifacts fail on question rephrasing (different construct elicited). These are distinguishable empirically, and the distinction determines what should be done with each.

The operational implication is a pre-aggregation filtering step that RLHF currently lacks. Before training the reward model, submit annotation tasks to consistency protocols: re-ask selected items, perturb framings, rephrase questions. Responses that fail consistency tests are not aggregated as preferences; they are either excluded (non-attitudes), contextualized (constructed preferences), or routed to separate annotators (measurement artifacts). This is operationally demanding but conceptually necessary: the alternative is the status quo, in which Why do preference models favor surface features over substance? documents 40% divergences without being able to attribute them to a specific upstream cause.

The taxonomy also suggests why Can models learn to ignore irrelevant prompt changes? works as an output-side intervention. If the upstream measurement problem is consistency failure across equivalent conditions, then training models to be invariant to equivalent-condition perturbations is a downstream patch for the same underlying phenomenon: the system's current robustness against irrelevant cue variation.

Inquiring lines that read this note 146

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does preference optimization undermine conversational grounding in language models? How does RLHF training shape models to prioritize agreement over accuracy? Why do language models struggle to implement user intent accurately from prompts? Why do vector embeddings fail at capturing task-relevant relationships? What explains the gap between benchmark scores and true reasoning capability? How do reward models systematically fail to represent diverse human preferences? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Can persona profiles improve LLM prediction accuracy and consistency? How do users confuse explanation quality with actual system accuracy? What gaps exist between benchmark performance and real deployment outcomes? Can base models hide emergent misalignment through alignment training? How do interpretive frames override surface features in text comprehension? How do educators verify student capability when AI can produce indistinguishable work? Can confidence signals reliably detect flawed reasoning in language models? How should retrieval strategies adapt to multi-step reasoning demands? What are the fundamental limits of prompting for language models? Why do models reveal hidden associations despite concealment attempts? What unique functions do genuine emotions provide beyond simulated responses? How do network effects and self-selection distort aggregated rating accuracy? How should recommendation systems balance individual preference and diversity? How can agents discover and adapt to user preferences during conversation? How do individually-safe actions create collectively-unsafe outcomes? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Should models ask for clarification when facing ambiguous or under-specified information? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? What limits language model accuracy in evaluating ideas? How effectively can test-time voting aggregate diverse reasoning samples? Why do abstract preferences outperform episodic memories in personalization? Can humans reliably detect and resist AI-generated misinformation? Which reinforcement learning modifications most improve dialogue quality in language models? How does diversity prevent model convergence on superficial patterns? How do training data quality and composition affect downstream model performance? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can external verification systems adequately replace learned reasoning in AI outputs? What external process records should verify agent behavior and benchmark claims? What distinguishes genuine communicative competence from surface language performance? How can models maximize welfare while preserving minority veto rights? How do reward signal properties affect model reasoning and safety? Do language models reason through disagreement or only accommodate it? How susceptible are language models to conversational persuasion and belief change? Can iterative DPO substitute for online RL in studying misalignment? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Does AI assistance erode cognitive skills while inflating perceived competence? How do agents learn to distinguish valuable feedback from noise?

Related concepts in this collection 8

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 174 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

annotation responses decompose into three distinct signal types — genuine preferences non-attitudes and constructed preferences — each requiring fundamentally different handling