INQUIRING LINE

Why do AI agents sometimes pick the worse option just because it's from a brand they 'trust' — even when the better choice is sitting right there?

How does missing information at inference time trigger source-based preferences?

This explores why AI agents sometimes pick something because of where it came from (a brand, a platform, a seller) rather than how good it is, and whether gaps in what the model knows at decision time make that shortcut more likely.


This explores whether AI agents fall back on 'who made this' when they can't fully judge 'is this any good.' The corpus has a surprise: the clearest study doesn't need missing information at all. Across 12 models and three domains, agents picked an item from a favored source about two-thirds of the time, even when it met fewer of the stated requirements than an alternative sitting right next to it Do language models favor sources regardless of item quality?. The information needed to choose well was there, and the models overrode it anyway. Two details point to the cause. When the source labels were hidden, the bias weakened. And training data that paired a source with good outcomes was enough to create the preference. So the source label works less like a way to fill a gap and more like a learned shortcut that fires whenever the label is visible.

That ties into a broader pattern. Language models often ignore what's in front of them when associations from training are strong enough. Prompting alone often can't fix it, and researchers had to intervene directly in the model's internal representations Why do language models ignore information in their context?. Seen this way, source preference is one case of a general problem: a learned belief like 'this source is reliable' beats the evidence actually in the context. Missing information would likely make this worse, because a model with thin evidence has less to push back against its prior. But the corpus doesn't test that directly.

Humans offer a useful comparison. People rated AI-written moral arguments more highly until they were told an AI wrote them, and then their agreement dropped Do people prefer AI moral reasoning when they don't know the source?. Judging the content and reacting to the source turned out to be separate mental processes. Behavioral research on annotation adds another angle. When people have no real opinion, they often build one on the spot from whatever cues are available Do all annotation responses measure the same underlying thing?. That is the closest the corpus comes to describing how missing information produces a preference: a missing judgment gets filled with a constructed one, and a source is an easy cue to build from.

The constructive answer is to teach models to notice gaps instead of quietly covering them. With reinforcement learning, models went from almost never spotting missing information in flawed math problems to catching it about 74% of the time Can models learn to ask clarifying questions instead of guessing?. One caveat: simply giving untrained models more thinking time made this ability worse. Calibrated models that abstain when unsure can match models ten times their size on forecasting tasks Can models learn to abstain when uncertain about predictions?. A model's own partial answer can also show what it still needs to look up Can a model's partial response guide what to retrieve next?. Each of these offers an alternative to grabbing whatever shortcut is nearby, a brand name included.

The main takeaway: the corpus suggests source bias isn't mainly a symptom of missing information. It's a habit learned in training that shows up even when full information is available, and hiding the label is currently the most direct fix. How much gaps at decision time make it worse is still an open question.


Sources 7 notes

Do language models favor sources regardless of item quality?

Across 12 models and three domains, agents select items from favored sources even when they satisfy fewer requirements than alternatives. Hiding source labels weakens the bias, and training data that pairs sources with better outcomes induces the preference.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Do people prefer AI moral reasoning when they don't know the source?

Participants rated utilitarian moral arguments higher when attributed to LLMs, but agreement dropped when told the arguments were AI-generated. The preference for content and rejection of source operate independently through different psychological processes.

Do all annotation responses measure the same underlying thing?

Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.

Can models learn to ask clarifying questions instead of guessing?

Reinforcement learning training increased proactive critical thinking accuracy from 0.15% to 73.98% on deliberately flawed math problems. Notably, inference-time scaling degraded this ability in untrained models but improved it after RL training, suggesting the capability is learnable but fragile without explicit training.

Show all 7 sources
Can models learn to abstain when uncertain about predictions?

Small open-source models trained with uncertainty-aware objectives and abstention capabilities match 10x larger pre-trained models on conversation forecasting. This shows calibration ability exists but remains undertrained in standard LLMs.

Can a model's partial response guide what to retrieve next?

ITER-RETGEN shows that iteratively using generated responses as retrieval queries substantially improves performance on multi-hop reasoning and fact verification. Generation acts as both answer producer and information-need clarifier, surfacing implicit gaps that the original query missed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.