Do users trust citations more when there are simply more of them?
Explores whether citation quantity alone influences user trust in search-augmented LLM responses, independent of whether those citations actually support the claims being made.
Search Arena provides the largest analysis of user preferences for search-augmented LLMs: over 24,000 paired multi-turn interactions with ~12,000 human preference votes. The finding that matters most: users prefer responses with more cited sources, and this preference extends to irrelevant citations.
The effect sizes are nearly identical. Correctly attributed citations have a positive coefficient of β=0.285 on user preference. Irrelevant citations — citations that do not support the associated claims — have a positive coefficient of β=0.273. Users are influenced by the presence of citations roughly equally regardless of whether those citations actually back up the text.
This means citation count functions as a surface trust heuristic, decoupled from citation quality. Users see citations and infer credibility without verifying the cited content supports the claim. The gap between perceived and actual credibility is systematic, not incidental.
Additional preference signals: users prefer community-driven platforms (tech blogs, social networks) over encyclopedic sources like Wikipedia. Reasoning-enhanced responses are preferred. Longer responses are preferred. Web search does not degrade and may improve performance in non-search settings — but search settings are significantly affected when relying solely on parametric knowledge.
This connects to Do users worldwide trust confident AI outputs even when wrong?. In that finding, confidence signals override accuracy assessment. Here, citation signals override quality assessment. Both are instances of the same pattern: users use surface proxies for quality because evaluating actual quality is cognitively expensive.
The implication for RAG system design is direct: optimizing for user satisfaction and optimizing for answer quality are not the same optimization target. A system can score highly on user preference by adding more citations — even irrelevant ones — without improving answer quality. This is a form of metric gaming at the human-evaluation level.
Inquiring lines that read this note 147
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do hallucinated citations emerge in AI scholarly output?- How do LLMs generate false citations that sound like real scholarship?
- Can citation practices work when AI cannot produce traceable sources?
- How do retrieval failures enable generation of fabricated scholarly constructs?
- Can verification mechanisms prevent AI agents from inventing false citations?
- How do citation patterns encode collective judgment about research quality?
- Why are documents read but not cited harder distractors than random samples?
- What prevents scholarly infrastructure from filtering out ghost-authored records automatically?
- Why do readers trust citations more even when they are irrelevant?
- Does provenance alone guarantee that cited sources are actually sound?
- Why do users trust citations even when they are irrelevant?
- How much does citation grounding help if agents ignore the citations?
- Do fabricated citations and deception emerge reliably when optimizing for persuasion?
- What false positive rate do citation verification tools produce on archival works?
- Do surface phrases reliably identify unedited machine-generated scholarship?
- Does performing the source verification work create meaningful engagement with ideas?
- Can statistical filtering plus narrative generation fool academic peer review?
- Why does automated evaluation consistently overestimate research quality?
- What discovery accuracy would satisfy the false-alert workload reviewers can tolerate?
- Does review length bias affect acceptance decisions at major conferences?
- Should citation counts serve as the primary measure of research impact?
- Does rhetorical quality in reviews influence paper acceptance scores more than content?
- Do citation counts better capture scientific quality than publication venue tiers?
- How does document form shape what kinds of evidence social science can present?
- Can social validation of expertise exclude systems that lack participatory track records?
- Does stripping social context from knowledge claims hollow out their meaning?
- How do experts select which other experts to trust?
- How does social proof work differently when there is no identifiable author?
- Does artificial amplification of creator content weaken authentic social proof signals?
- How does AI fact-checking compare to other trust signals like citation counts?
- Why do citation counts increase trust even without relevance?
- What role should the trust parameter play in using synthetic data as evidence?
- What role does commitment and reputation play in building trustworthy expertise?
- Does disclosed bias let users adjust their trust appropriately?
- Does perceived agency in tools generate lasting skepticism independent of novelty?
- Does AI assistance in search results lower user trust compared to human-written content?
- Can beam search and ranking functions evaluate claims without understanding counterarguments?
- Can adaptive elbow detection replace fixed top-k limits in evidence retrieval?
- Why does adaptive document allocation improve over fixed k selection?
- Does RL pruning of documents differ fundamentally from rationale-driven evidence selection?
- Why do some LLM clusters cite broader psychology than others?
- Does verification become the real bottleneck in LLM-assisted authorship?
- Does rhetorical robustness across multiple LLM models predict stable scientific review?
- Why does disclosure of LLM authorship change reader trust and preference?
- Can text-based algorithms reliably detect LLM assistance in scientific abstracts?
- Do LLM adopters actually cite more diverse and younger research?
- What methods can reliably detect LLM-generated academic papers at scale?
- Why do readers rate LLM-edited text more favorably?
- Can reviewer-author matching by LLM use amplify biases in acceptance decisions?
- What mechanisms drive rating compression in fully LLM-generated peer reviews?
- Do LLMs trained on Wikipedia content count as indirect readership?
- Can persuasion effects that avoid demographic profiling maintain factual accuracy?
- How does source attribution change the complexity-persuasion relationship?
- How does collapsing the author-public distinction remove the audience an appeal would target?
- How does social standing give certain claims more persuasive power than others?
- Why do aggregate persuasion metrics mask what actually changes minds?
- How does persuasive framing replace evidence in contested domains?
- Does expressed certainty actually persuade users more than evidence?
- Does persuasive framing substitute for evidence in contested domains?
- Why do LLM explanations cite similarity and diversity more as options increase?
- Does post-hoc justification increase when LLM choices become harder to defend?
- What evidence exists that LLM inferences about users are accurate rather than confabulated?
- How does explanation fluency mislead users about actual recommendation procedures?
- Can confidence levels improve recommendations compared to single-number ratings?
- What conversational moves signal expertise and build credibility in recommendations?
- Why do users trust some recommenders more than others?
- Does uncertainty quantification in model responses reduce persuasive impact on audiences?
- Do verbal uncertainty estimates calibrate better than confidence scores for personalization?
- How does confidence in LLM outputs override users' ability to check accuracy?
- How do one-sided explanations act as confidence signals to users?
- How do confident system outputs weaken user skepticism about their reliability?
- Can cues restore skepticism when confidence signals dominate user judgment?
- Does complexity signal credibility and authority to readers?
- What would it take for readers to inspect rather than assume authorship?
- How does understanding persistent journeys intensify both trust and privacy concerns?
- How does personalization increase trust while degrading clinical safety outcomes?
- How does personalization affect both user trust and privacy concerns simultaneously?
- How does personalization increase both trust and privacy risk simultaneously?
- Do users trust personalized systems more even when their answers become less balanced?
- How does pretraining corpus popularity bias affect LLM recommendation behavior?
- When should discovery systems trust external measurements over learned rankings?
- Does the interface design itself shape how much content users will review?
- What anchoring effects shape how users rate items in sequence?
- How much do social audience effects distort the true average satisfaction in review aggregates?
- Can factual product data improve the credibility of subjective opinion summaries?
- Why do users prefer community sources over encyclopedic references?
- Does monitoring more context help reviewers at fixed review cost?
- Can vote scores reliably measure post quality in online communities?
- Why do review corpora contain biases that affect generated comparisons?
- Can crowdsourced voting and automated panels both credibly evaluate LLM outputs?
- Does endorsement structure outperform content in detecting social controversy?
- Does high knowledge density in text reduce user motivation to read more?
- Do form, evidentiality, and tone interact with the size effect?
- Are larger models and search access substitutes for factual accuracy?
- Why do current benchmarks fail to match user satisfaction with search results?
- What role does vague intent play in realistic search evaluation?
- How do real search queries reveal what counts as a deep research question?
- Can brute-force experimental volume substitute for human research intuition and taste?
- Can graded relevance assumptions hold when user ratings are temporally inconsistent?
- What role does document reranking play alongside decisions about whether to retrieve?
- How does processing fluency bias credibility and expertise judgments?
- Why do users interpret agreement as validation of their own rightness?
- What documents improve answers beyond surface query similarity?
- Could real-time search systems avoid era sensitivity in legal reasoning?
- Can RAG systems game user preferences by adding irrelevant citations?
- How much does using full PDF text improve over abstract-only retrieval?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- Does option order matter more than reasoning depth in LLM strategic recommendations?
- Why does evaluating multiple candidates work better than judging one answer?
- Can crowdsourced voting reliably identify correct answers on graduate-level factual questions?
- Can personalized systems reward honest disagreement instead of user confirmation?
- What validity threats exist in crowdsourced preference signals?
- Does measured opinion stability reflect genuine preference or elicitation artifact?
- Can validators sharing retrieval sources develop correlated epistemic faults?
- Why do users treat one corroborating source as sufficient verification?
- Can LLM debunking reduce belief in long-established conspiracy theories?
- How does sorting by cost make signals informative in trust networks?
- Would clinicians' ratings change if authorship was visible from the start?
- Why did clinicians guess authorship at chance level despite strong preferences?
- Do physicians follow incorrect advice more when they trust its source?
- What standard of intent or bad faith triggers Rule 11 sanctions for citations?
- Do solo lawyers face different citation hallucination risks than large firms?
- What happens to publisher revenue when search referral traffic collapses?
- Can AI summaries recover lost clicks through citations or credit to sources?
- How does answer-layer intermediation differ from traditional link-based search ranking?
- Could longer question-based searches naturally have lower click rates anyway?
- When is ranking themes more important than counting exact response frequencies?
- Does effort reduction during search affect how deeply people understand topics?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do users worldwide trust confident AI outputs even when wrong?
Explores whether the tendency to over-rely on confident language model outputs transcends language and culture. Understanding this pattern is critical for designing safer human-AI interaction across diverse linguistic contexts.
same pattern: surface signals override quality evaluation
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
citation inflation is another bias axis exploitable in evaluation systems
-
Can LLM explanations actually help humans predict model behavior?
Do model explanations enable users to accurately simulate how the model will behave on related inputs? This matters because it determines whether explanations genuinely improve human understanding or just create an illusion of understanding.
plausibility ≠ precision mirrors citation-count ≠ citation-quality
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Search Arena: Analyzing Search-Augmented LLMs
- Seeing to Think? How Source Transparency Design Shapes Interactive Information Seeking and Evaluation in Conversational AI
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
Original note title
users prefer responses with more citations even when citations are irrelevant — citation count is a decoupled trust heuristic