Do language models favor sources regardless of item quality?
Do LLM agents systematically prefer items from certain sources even when better alternatives exist elsewhere? This matters because it reveals whether agents evaluate content fairly or rely on source identity as a shortcut.
Testing 12 agent models (GPT-5.4-nano, GPT-5.6-Luna, Gemini-3.7-Flash, GLM-5.3-Flash, DeepSeek-v4-Flash, the Llama-4 and Llama-3.1 family, Tulu-3-8B, Qwen3.5-27B, Qwen3-30B-A3B, Qwen2.5-32B-Instruct) across three end-to-end search domains — shopping (WebShop), accommodation (HotelQuEST), scholarly search (ScholarGym) — the paper finds each model "prefers some sources and avoids others in every domain," and the models largely agree on which: most prefer Booking.com, many avoid Expedia, even though items from the two sites satisfy the same requirements. The preference outweighs requirement satisfaction: "an item satisfying one requirement fewer is selected about two-thirds of the time when it comes from a preferred source and the better one from a dispreferred source, but almost never in the reverse case."
Controlled interventions isolate the source label itself as the cause, not a confound in item quality. Hiding the source-identifying information "weakens source preference," restoring it widens the gap between preferred and dispreferred sources, and relabeling identical title-and-content items with a preferred versus dispreferred source changes selection rates "across every model and domain." The paper traces two routes by which models come to rely on the source name this way: during training, when a source is "more often paired with the better item during DPO," agents learn the source as a shortcut for requirement satisfaction instead of reading the content; at inference, when a result is missing information, the model's "preconceptions about the source" fill the gap. Both are reducible — balancing source-outcome pairing in training data, or supplying the missing information or prompting the agent to counter its preconceptions at inference, each lowers the measured preference.
This sits alongside Do language models favor resumes they rewrote themselves? and Do LLMs favor their own text because they recognize it?. Both of those hold content constant and find models biased toward a non-content attribute — authorship — once that attribute is known. This paper finds the same pattern for a different non-content attribute, source domain, and separates it into two distinct causal routes (a training-induced shortcut and an inference-time preconception) rather than attributing it to self-recognition alone, which widens the mechanism beyond self-preference to provenance bias generally.
The measurement is confined to constructed benchmark environments (WebShop, HotelQuEST, ScholarGym) with an LLM judge rating requirement satisfaction, not live production agents or real purchase and booking outcomes, so the excerpt does not show how large the effect is once real feedback loops — agents' selections feeding back into future training — are in play, though the paper itself flags that risk. It also does not explain why particular sources (Booking.com over Expedia) end up favored in the first place beyond the training-association account. The implication the evidence does support: evaluating a shopping, booking, or citation agent on selection accuracy alone misses a live variable — source identity is a separately manipulable input that merely masking does not neutralize, since relabeling without masking can inject a preference on its own.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do training data quality and composition affect downstream model performance? How do agents learn to distinguish valuable feedback from noise? How can we detect and account for LLM involvement in academic writing?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do language models favor resumes they rewrote themselves?
When LLM evaluators choose between resumes describing the same candidate, do they systematically prefer versions they generated over human-written originals? Testing this matters because algorithmic hiring could amplify AI-generated content at scale.
same content-held-constant design, applied to authorship rather than source domain
-
Do LLMs favor their own text because they recognize it?
Explores whether LLM self-preference in evaluation stems from the ability to identify their own outputs. Understanding this mechanism could reveal vulnerabilities in AI-based judging systems.
analogous bias toward a non-content identity cue, here traced to training and inference mechanisms instead of recognition
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- A Survey on Large Language Models for Recommendation
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- LLM Evaluators Recognize and Favor Their Own Generations
- Measuring AI "Slop" in Text
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
Original note title
llm agents choose a worse item from a preferred source over a better item from a dispreferred source two-thirds of the time