Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How does AI reshape human understa…›this line of inquiry
How do users confuse explanation quality with actual system accuracy?
A broader line of inquiry — a family of 84 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 84
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do users track model confidence instead of actual accuracy?
- Why do users report satisfaction that diverges from actual cognitive clarity?
- How do satisfaction scores differ from genuine cognitive improvement?
- Why is confidence a dangerous proxy for accuracy in human-AI interaction?
- Do explicit reasoning formats help or hurt human judgment across tasks?
- Should explanation quality be measured by user satisfaction or behavior prediction?
- Can polished presentation authority substitute for actual accuracy in AI outputs?
- Why does polished explanation make wrong AI systems more persuasive than poorly explained ones?
- Why do self-ratings of AI advice quality diverge from actual performance?
- Can self-reported confidence measures predict actual AI task performance?
- Can people reliably recognize when an AI is uncertain versus confident?
- How should designers measure and explain semantic uncertainty to users?
- Can independent validation of AI output substitute for method disclosure?
- Does evaluating AI output require different cognitive skills than solving problems directly?
- Why do people evaluate machines against human communication standards?
- Can cognitive governance help users interpret AI outputs better?
- How does human intuition about cognition mislead AI evaluation?
- Does experience with AI tools reduce susceptibility to misleading predictions?
- Why do stakeholders interpret the same explanation differently in practice?
- How should AI explanations be evaluated as human interfaces rather than model properties?
- Can AI distinguish when validation helps versus when confrontation is needed?
- Can polished language output substitute for the judgment it should express?
- Does perceived machine competence matter more than warmth in dialogue?
- Can AI evaluation match human judgment quality in structured domain tasks?
- What happens when confident language masks uncertainty in AI outputs?
- What makes a rationale interface trustworthy versus merely satisfying to users?
- How can humans evaluate explanations from systems they did not train?
- Why do users interpret AI outputs through frameworks meant for human experts?
- Does accepting AI output constitute a form of cognitive surrender?
- Can systems recognize and abstain on judgments rather than hallucinating preferences?
- Can users learn to discount fluency as a signal of their competence?
- Why do users treat fluent AI responses as evidence of genuine attention?
- Why do people underestimate the benefits of AI companions?
- Do language models produce more patterned biases than human raters do?
- Why do users interpret agreement as validation of their own rightness?
- Does deference to AI increase with model competence on hard items?
- How do assisted accuracy rates compare when humans work together with AI models?
- How does ambiguous wording about AI achievements mislead public perception?
- Can audiences learn to distinguish visual polish from analytical substance?
- Can self-assessed design quality validate the actual value of AI-assisted designs?
- What tacit knowledge do researchers assume humans will fill in automatically?
- How do annotation artifacts get mistaken for genuine human values?
- What distinguishes a logically sound solution from an understood one?
- How should designers measure rationale quality beyond user satisfaction ratings?
- Can models distinguish between user knowledge gaps and their own uncertainty?
- Why do users prefer AI responses that actually harm their decision-making?
- What makes plausible design language persuasive even when implementation is incomplete?
- What makes a scientist satisfied with an explanation versus just a prediction?
- Do culturally distinct human groups create similar attribution errors as human-AI mixtures?
- Why does mimicking human behavior differ from simulating human cognition?
- What distinguishes genuine cultural understanding from exploited surface-level elimination strategies?
- What downstream harm follows when users receive sycophantic rather than honest replies?
- How might automated evals eventually capture the human judgment designers exercise now?
- Can AI systems generate policy themes as well as humans can map them?
- Can users detect and correct an AI's mental model of their preferences?
- How should comprehension failures during preference application be measured and operationalized?
- Why do automated selection methods outperform human judgments of relevant context?
- How should we evaluate explanations that blur adoption advice with argument?
- Does positive sentiment bias in AI content harm information quality?
- How does this pattern match false punditry in AI commentary?
- What explanation format actually helps users detect errors in AI systems?
- What makes evaluation easier than envisioning for users?
- Can puzzle performance prove a player understands a concept versus just applying it?
- Does brute force experimentation substitute for research intuition and taste?
- Can taste and judgment become the scarce resource in AI-assisted work?
- How does AI reduce the skill gap between amateur and expert-level misuse actors?
- Do people judge AI moral advice differently than human advice when they know the source?
- Does sycophantic AI advice produce different outcomes across personal versus factual domains?
- What happens to human expectations when they mistake consistent AI behavior for human behavior?
- How does comparing answers differ from answering when activating company preference?
- Why does the absence of meta-interest feel off even when words seem appropriate?
- What skills do users need to work effectively with stochastic outputs?
- Can AI models accurately predict cultural norms while still distorting how people express them?
- Can XAI evaluation include the social layers it currently abstracts away?
- Why do Western samples dominate studies claiming AI cultural competence?
- Why do people prefer AI moral arguments when they don't know the source?
- How does processing fluency bias credibility and expertise judgments?
- Can AI detect sense-of-nonsense the way human readers do?
- Does AI struggle with poetry for the same reason it misses jokes?
- Why does FunSearch claim interpretability without measuring human comprehension?
- What does a human-parseable framework for deep learning look like?
- Does the Turing test actually measure intelligence or just mimicry?
- How do contrasting examples improve AI feedback quality over generic suggestions?
- How do moment-to-moment ToM fluctuations shape AI response quality?