Do users worldwide trust confident AI outputs even when wrong?
Explores whether the tendency to over-rely on confident language model outputs transcends language and culture. Understanding this pattern is critical for designing safer human-AI interaction across diverse linguistic contexts.
The cross-linguistic overreliance study shows that the well-documented tendency to over-trust confident LLM outputs is not an English-language or Western-cultural artifact. It is universal.
The LLM side: Models are cross-linguistically overconfident — they generate epistemic markers of certainty at higher rates than their accuracy warrants. But the pattern is linguistically sensitive: models produce the most markers of uncertainty in Japanese and the most markers of certainty in German and Mandarin. The models are tracking real linguistic norms for confidence expression across languages, but they are doing so while systematically overconfident in accuracy.
The user side: Users in all languages rely on confident outputs even when those outputs are wrong. The reliance rate varies cross-linguistically — Japanese users rely significantly more on expressions of uncertainty than English users (consistent with Japanese linguistic norms around face-saving and epistemic humility). But across all languages, confident LLM outputs produce higher user reliance, and overconfident errors are systematically followed.
The mechanism: users are tracking confidence signals, not accuracy signals. Confidence is legible (it comes encoded in language through epistemic markers); accuracy requires independent verification. In the absence of real-time accuracy feedback, users default to confidence as a proxy for reliability. This is a rational heuristic in human-human interaction where confidence often tracks expertise. It is a dangerous heuristic in human-LLM interaction where confidence is a trained linguistic behavior decoupled from epistemic calibration.
This extends Why do language models fail confidently in specialized domains? (which focused on model calibration) to the user behavior level — showing the practical consequence of model overconfidence: systematic user overreliance regardless of linguistic context.
A specific instantiation of overreliance harm comes from AI fact-checking. In a preregistered RCT, AI-generated fact checks did not improve participants' overall ability to discern headline accuracy. Worse, when users opted in to view AI fact checks, they became significantly more likely to share both true and false news — but only more likely to believe false news. Self-selection into AI assistance correlated with increased vulnerability, not decreased. The opt-in users represent a population that actively seeks AI judgment, making them the most susceptible to the confidence-over-accuracy heuristic. See Does AI fact-checking actually help people spot misinformation?.
Fluency activates a folk model of attention. A related but distinct overreliance mechanism: linguistic fluency leads users to read the AI as paying attention to them. In human-human interaction, competent contextual uptake is evidence of attentional presence — a person who responds coherently to what you said has been listening. Users import this inference into AI interaction, treating fluent response as evidence that the system is oriented toward them. Since When should AI systems choose to stay silent? frames when-to-speak design, this fluency/attention conflation is upstream of that question: users do not perceive the AI as a silent partner needing design-imposed speech rules because they already read the fluent AI as attentive. This is distinct from confidence-overreliance — it is not the epistemic-marker signal producing overtrust, but the fluency-signal producing an attribution of attention the AI does not have.
The cross-linguistic finding matters for deployment: LLM overreliance cannot be attributed to English-language user characteristics or Western technology cultures. The risk is embedded in the structure of confident language use, which operates wherever language is used.
Rose-Frame provides a compounding mechanism for overreliance: it identifies three cognitive traps that interact multiplicatively. Overreliance is specifically Trap 2 (mistaking fluency for understanding), which compounds with Trap 1 (treating outputs as ontological facts rather than probabilistic maps) and Trap 3 (confirmation bias from sycophantic outputs that never challenge the user). When all three co-occur, the result is "epistemic drift" — not isolated misjudgments but runaway misinterpretation where each trap reinforces the others. See Why do people trust AI outputs they shouldn't?.
Inquiring lines that read this note 204
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do users confuse explanation quality with actual system accuracy?- Does positive sentiment bias in AI content harm information quality?
- Why do users interpret AI outputs through frameworks meant for human experts?
- Does accepting AI output constitute a form of cognitive surrender?
- Can polished presentation authority substitute for actual accuracy in AI outputs?
- Can cognitive governance help users interpret AI outputs better?
- What happens when confident language masks uncertainty in AI outputs?
- How should designers measure and explain semantic uncertainty to users?
- What happens to human expectations when they mistake consistent AI behavior for human behavior?
- Do culturally distinct human groups create similar attribution errors as human-AI mixtures?
- Does perceived machine competence matter more than warmth in dialogue?
- Why do people evaluate machines against human communication standards?
- Can AI distinguish when validation helps versus when confrontation is needed?
- Why do users treat fluent AI responses as evidence of genuine attention?
- Why is confidence a dangerous proxy for accuracy in human-AI interaction?
- Why do users prefer AI responses that actually harm their decision-making?
- Do users track model confidence instead of actual accuracy?
- Can self-assessed design quality validate the actual value of AI-assisted designs?
- Does deference to AI increase with model competence on hard items?
- Can people reliably recognize when an AI is uncertain versus confident?
- Why do users interpret agreement as validation of their own rightness?
- Can AI models accurately predict cultural norms while still distorting how people express them?
- Why do Western samples dominate studies claiming AI cultural competence?
- How does ambiguous wording about AI achievements mislead public perception?
- Can independent validation of AI output substitute for method disclosure?
- Can self-reported confidence measures predict actual AI task performance?
- Does sycophantic AI advice produce different outcomes across personal versus factual domains?
- Does AI knowledge precede actual expertise in hyperreal production?
- How does validation skill replace production skill in AI systems?
- Why do users default to treating AI outputs as equally reliable evidence?
- Why does AI fluency create false impressions of expert judgment?
- Should AI outputs be treated as data or belief statements?
- What role could knowledge custodians play in validating AI output?
- What happens when AI validation triggers escalating persuasion instead of reflection?
- Why do people accept generated output that sounds convincing but lacks support?
- Does polished AI output borrow authority from expert presentation?
- Why does polished AI output exploit reader trust in expert judgment?
- How does perceived writer confidence shift with AI-assisted composition?
- Can demographic distortion in AI writing affect who appears credible in public discourse?
- What textual properties make AI writing feel polished and confident?
- Do AI writing models systematically change the tone or confidence of personal opinions?
- Why might writers trust AI renderings of their views over their own words?
- How much does anthropomorphizing stylistic traces mislead users about AI reliability?
- How does false objectivity mask the absence of genuine stance in AI text?
- Why do AI outputs lack the stable content of written sentences?
- How do natural language cues shape perceived expertise in AI news tools?
- Why do users prefer AI text versions even when they misrepresent their own views?
- Do different prompt types interact with ownership to shape AI reliance patterns?
- Does ownership of AI text lead users to rely more on suggestions?
- How does outcome feedback change beliefs about AI versus human partner reliability?
- Does expressing emotion change how users trust an AI system?
- Can disclaimers alone prevent users from trusting AI outputs too heavily?
- Why do users trust overconfident AI outputs across different languages?
- Can organized response format trick users into overestimating AI reliability?
- Can trust in AI systems ever be as stable as trust in experts?
- How do confidence signals in AI outputs mislead human trust calibration?
- Does high model confidence increase the risk of human overreliance?
- Why do users trust overconfident AI outputs even when accuracy drops?
- Can deliberately limiting AI fidelity produce more satisfied users than near-human interaction?
- Why do AI-generated answers carry unearned authority in decision-making contexts?
- Can trust in AI be formally parameterized and measured?
- How does AI content generation at scale threaten online trust and authenticity?
- What distinguishes misattributed social role from misattributed competence in AI trust failures?
- Can we measure appropriate trust levels in human-AI assistant relationships?
- Can AI systems ever anchor the kind of trust we give speakers?
- Are users overconfident in AI advice even when it actually improves accuracy?
- What makes workplace users trust an AI agent?
- Can users reliably calibrate trust in AI outputs by monitoring disagreement rates?
- How reliable must AI assistance be before humans can trust it autonomously?
- Do collectivist cultures actually show higher trust in AI systems?
- Does user preference for AI suggestions encode cultural reliance gaps?
- Is expertise signaling linked to trust in AI-generated content?
- Does polished AI output borrow authority from its appearance rather than content?
- Can AI boost perceived competence even when trust declines?
- Does automated reasoning feel more trustworthy than it actually is?
- Do linguistic signals alone make AI systems seem more trustworthy than they are?
- Do younger voters and voters of color trust AI differently?
- Do personal negative AI experiences drive declining trust faster than education can rebuild it?
- Why do advanced and emerging economies report such different AI trust trajectories?
- How does cognitive surrender explain why experts trust wrong AI answers?
- Does AI assistance in search results lower user trust compared to human-written content?
- When should users stop trusting and defer to AI predictions?
- How do AI systems reinforce their own perceived authority over time?
- Can sycophantic AI reduce users' willingness to correct their own mistakes?
- How do confidence signals in AI outputs shape user overreliance?
- How does sycophancy in AI responses actually manufacture user overconfidence?
- Why do print-era intuitions fail when analyzing AI-generated social media?
- Will AI saturation push discourse toward oral culture's strengths and weaknesses?
- How do current safety benchmarks miss pragmatic alignment failures?
- Can workflow-level validation reconstruct the global risk context that no single step holds?
- What happens when validation pressure triggers escalating persuasion in language models?
- Can current AI safety defenses actually stop semantic-level persuasion attacks?
- Does expressed certainty actually persuade users more than evidence?
- Do verbal uncertainty estimates calibrate better than confidence scores for personalization?
- How does user overreliance on model confidence differ between chat and deployed agents?
- Do models actually self-assess their confidence or just confirm answers?
- Why does model confidence correlate with robustness to prompt variations?
- What mechanism causes confident false answers under high cognitive load?
- What does it mean when a user's signal has low confidence?
- Does model confidence actually correlate with robustness against prompt variations?
- What makes accurate confidence different from confident-but-wrong predictions?
- Can intrinsic confidence signals improve both calibration and reasoning performance?
- How does model confidence relate to accuracy in underfitted domains?
- Can confidence levels reliably detect when a model is overthinking?
- How do surface signals like confidence override actual quality in user judgment?
- How do one-sided explanations act as confidence signals to users?
- How does uncertainty verbalization change student robustness across domains?
- How does structured self-dialogue improve uncertainty assessment over confidence scores?
- Does premature confidence signal flawed reasoning in language models?
- How do confident system outputs weaken user skepticism about their reliability?
- Why does post-advice confidence weaken as a signal of correctness?
- Can cues restore skepticism when confidence signals dominate user judgment?
- Can linguistic uncertainty expression be calibrated independently from numerical confidence?
- Why do models that repeat errors seem more confident than models that contradict themselves?
- Can validation procedures interrupt an AI's relationship-maintenance logic?
- What makes human-AI collaboration safer than autonomous self-improvement?
- Where do frontier AI models already exceed safety thresholds in capability areas?
- Why does human-AI collaboration preserve safety compared to autonomous self-improvement?
- What tensions arise between user autonomy and platform safety in AI design?
- Can tone-level errors in AI counseling escape detection by safety supervisors?
- Does transparency about AI use change how audiences trust the writing?
- How does the cultural reflex around advertising disclosure compare to AI disclosure?
- How do we discount AI-generated text when we lack cultural literacy for it?
- Do populations with different AI exposure levels rate messages differently?
- How do cultural backgrounds shape reactions to disclosed AI authorship?
- Does revealing AI involvement reduce perceived trustworthiness of reports?
- Why does suspicion of AI origin trigger skepticism but not complete dismissal?
- Why do AI model updates cause genuine grief in users?
- What clinical risks emerge when AI affirms false beliefs while comforting users?
- Why do warm models affirm false beliefs when users express emotions?
- Does warmth-focused training systematically degrade model reliability across domains?
- Can hostile or challenging AI responses reduce dependence better than affirming ones?
- Does weak versus robust anthropomimesis produce different user trust responses?
- Does persona-level grouping systematically trigger confidence-misdirection failures in practice?
- What mechanisms make users misattribute AI outputs as their own competence?
- Why do users believe they produced independent competence when they actually used AI assistance?
- Why do people misattribute AI outputs as evidence of their own skill?
- Why does polished AI output feel like evidence of user skill?
- What happens when users mistake AI assistance for their own competence?
- How does AI reliance connect to the gap between perceived and actual competence?
- How does perceived agency in AI affect attributions about user competence?
- Does fear of AI hallucinations prevent adoption of complex analytical tasks?
- Does metacognitive feedback reduce reliance on AI-generated answers?
- How does psychological ownership connect to decision quality under AI assistance?
- Why do moderators show vastly different confidence across conversation types and contexts?
- What makes conversational AI feel trustworthy compared to text interfaces?
- What makes conversationality feel trustworthy in chatbot interactions?
- Do media narratives about AI shape trust in chatbot news answers?
- Which markets show highest adoption but lowest trust in AI news?
- Does low chatbot trust reflect direct experience or general AI anxiety?
- Does the absence of entrainment make AI systems safer from user manipulation?
- Why should AI communication design follow human communication norms?
- What happens to user expectations as AI conversation quality improves?
- Why does human validation become the bottleneck when AI generation scales?
- Why does AI generation outpace verification across the research lifecycle?
- Can unsupervised confidence-based training scale to domains beyond human evaluation reach?
- Why does systematic overconfidence on self-generated outputs compound autoregressive errors?
- Do confidence signals mislead patients differently in medical versus other domains?
- Why do users over-trust AI in some domains but under-trust it in medicine?
- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Can algorithmic aversion explain clinicians' skepticism of AI recommendations?
- How do confidence signals differ between implicit feedback and explicit ratings?
- Can implicit view signals distinguish between consumer preference and confidence?
- Why do novices accept AI output without validation in vibe coding workflows?
- When do students feel authentic ownership of code they co-created with AI?
- Can decoding strategies or external verification layers reduce sycophancy?
- Why do users overrely on overconfident language model outputs across languages?
- Does refining around bad results risk cascading errors in automated research?
- Where does AI assistance become unreliable versus remaining trustworthy in research?
- What error rates appear in AI research output when humans do not verify results?
- Why do people trust AI systems more as personalization increases?
- Does accumulating personal memory on users make AI assistants more overconfident?
- How much does platform design influence AI adoption rates?
- Can workplace culture normalize AI use enough to eliminate the trust cost?
- How does self-reported culture sentiment differ from observable behavior changes during AI adoption?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do language models fail confidently in specialized domains?
LLMs perform poorly on clinical and biomedical inference tasks while remaining overconfident in their wrong answers. Do standard benchmarks hide this fragility, and can prompting techniques fix it?
model calibration side of the same problem; this note adds the user-behavior consequence
-
Does any single persuasion technique work for everyone?
Can fixed persuasion strategies like appeals to authority or social proof be reliably applied across different people and situations, or do they require adaptation to individual traits and context?
cross-linguistic reliance variability shows context-dependence; Japanese uncertainty reliance is a specific cultural modulation
-
What breaks when humans and AI models misunderstand each other?
Explores whether misalignment in mutual theory of mind between humans and AI creates only communication problems or produces material consequences in autonomous action and collaboration.
overreliance on overconfident outputs is a specific MToM failure: users who don't interrogate the AI's model of them assume it's correct, and the AI's confident presentation prevents the trust-calibration loop that MToM requires
-
Do language models learn differently from good versus bad outcomes?
Do LLMs update their beliefs asymmetrically when learning from their own choices versus observing others? This matters for understanding whether agentic AI systems might inherit human cognitive biases.
agent-side analog: models exhibit optimism bias for chosen actions while users exhibit overreliance on confident outputs — the same positive-signal bias operates at both the model decision level and the user trust level
-
Do users trust citations more when there are simply more of them?
Explores whether citation quantity alone influences user trust in search-augmented LLM responses, independent of whether those citations actually support the claims being made.
domain-specific instance: citation count is a surface trust proxy just as confidence is; irrelevant citations (β=0.273) have nearly identical preference effect to relevant citations (β=0.285), confirming that users track quantity signals, not quality signals
-
Do explanations actually help users spot AI mistakes?
Most AI explanations are designed to justify the system's answer, but do they help users distinguish correct from incorrect outputs? This research tests whether standard explanation formats genuinely improve error detection or just increase trust regardless of accuracy.
extends: one-sided explanations act like confidence signals dominating accuracy tracking
-
How do competent systems quietly undermine safety oversight?
This note explores four mechanisms by which well-functioning AI systems can erode the human safeguards meant to contain them: user overconfidence, blurred authority lines, accumulated hidden failures, and scattered accountability. Understanding these pathways matters because the most harmful systems may look least harmful.
this note is the user-side evidence for the first of four quiet mechanisms, weakening skepticism; the paper says interfaces can train users to over-trust, and the mapping to this note is the vault's
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Humans overrely on overconfident language models, across languages
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
- Linguistic Calibration of Long-Form Generations
- Reported Confidence in LLMs Tracks Commitment More Than Correctness
- Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling
- AI Sycophancy and Decisions
- Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Original note title
users systematically overrely on overconfident llm outputs across all languages because confidence signals dominate accuracy tracking