The Decision to Verify: How Warmth and User Characteristics Shape Reliance on Conversational Agents for Information Search

Paper · arXiv 2605.28498 · Published May 27, 2026
Knowledge After the Web

Abstract Conversational artificial intelligence (AI) provides an efficient and convenient gateway to information access. However, it can cause overreliance when users blindly trust AI and accept its answers without fact-checking. Information search increasingly follows a hybrid interaction paradigm that combines conversational AI with web search, making fact-checking easier. In this paper, we examine whether this interaction paradigm is effective in curbing reliance. We further investigate the underlying factors (e.g., digital literacy and conversation warmth) that drive users to verify AI answers. We conduct a mixed-subjects question-answering experiment where participants interact with either a warm or a neutral chatbot. Our findings reveal that reliance persists despite users having access to both conversational and web search. The decision to verify is driven primarily by existing user perceptions (e.g., prior trust in chatbots) rather than answer properties, with some users factchecking regardless of the context and others trusting chatbots by default. Warm conversational style has an indirect yet critical influence on reliance by increasing agreement with the chatbot when it is incorrect. Consulting additional AI sources predicts higher accuracy, while traditional web search does not. Our study extends overreliance research by: (a) demonstrating its persistence despite access to fact-checking, (b) identifying verification behavior as user-dependent, and (c) revealing conversational warmth’s indirect effect

Introduction. Information search has seen a rapid change in recent years, becoming more conversational as the popularity of Large Language Models (LLMs) has grown (Chatterji et al., 2025). Approximately 10% of the world’s adult population uses ChatGPT (OpenAI, 2022) at least weekly, with one in four queries involving information seeking (Chatterji et al., 2025). Compared to traditional search interfaces, such as search engines, conversational agents offer better accessibility, are easier to use, and require less effort (Liang et al., 2025; Spatharioti et al., 2025; Kaiser et al., 2025). Despite the advantages, there exists a major downside of conversational agents: hallucinations. Humans rely on AI outputs even when they are incorrect, showing overreliance (Klingbeil et al., 2024). Overreliance leads people to accept the first answer they receive from an LLM without further inquiries in challenging tasks (Spatharioti et al., 2025; Kim et al., 2024). Furthermore, chatbots give an inflated sense of trust and confidence, even when they perform worse than search engines (Xu et al., 2023; Mayerhofer et al., 2025). LLMs hallucinate responses that seem plausible at first glance, albeit being nonfactual (Huang et al., 2025), making overreliance a critical issue for conversational agents (Xu et al., 2023; Spatharioti et al., 2025). In controlled experiments, participants are generally constrained to interact with either a chatbot or with a search engine (Spatharioti et al., 2025; Liang et al., 2025). However, in reality, users can access both: Recent developments have introduced a new search paradigm that provides users with simultaneous access to conversational agents (e.g., chatbots) and web search (Reid, 2024). Due to this new paradigm, hallucinations are not regarded as a major issue for most users, since chatbot responses can be easily verified by external sources (Skjuve et al., 2024; Menon and Shilpa, 2023). It can be argued that the new paradigm solves overreliance by combining the strengths of both conversational and traditional approaches. However, there is evidence showing that the mere presence of AI can induce reliance (Kim et al., 2024). Participants seek external confirmation less frequently when interacting with AI and tend to agree more with AI responses regardless of accuracy (Kim et al., 2024). Moreover, chatbots still hallucinate and give incorrect answers, even with web search (NewsGuard Technologies, 2025). These findings suggest that the new search paradigm might not be the definitive answer against overreliance. Despite verification tools being readily available with the new paradigm, the decision to fact-check or trust an AI answer is up to the user. Almost 90% of the utterances contain check-worthy claims in information-seeking tasks (Joko et al., 2025), making it unrealistic for users to verify all responses. Users have to make a conscious decision about which answers need further verification. However, users fail to fact-check results that confirm existing beliefs (Elsweiler et al., 2025) or become anchored to the first information source they encounter (Aslett et al., 2024). Consequently, tool availability alone cannot ensure a solution to reliance, as the ultimate barrier remains the human decision-making process itself. Previous experience, AI expertise, and digital literacy shape fact-checking behavior by reflecting users’ awareness of LLM limitations (Skjuve et al., 2024; Wang et al., 2024). Similarly, trust in AI systems determines whether users accept AI outputs as is (Klingbeil et al., 2024; He et al., 2025a). Beyond user background, chatbot conversational style, especially its warmth, influences reliance by shaping risk perceptions and emotional trust (Xiao et al., 2024; Jung et al., 2024; Takayanagi et al., 2025). In summary, overreliance appears to persist despite the emerging trend of integrating conversational AI with traditional search. Existing studies cannot conclusively explain why and how overreliance manifests when users have simultaneous access to both modalities (Spatharioti et al., 2025; Liang et al., 2025). Addressing this gap requires a deeper understanding of the psychological and contextual factors that shape users’ verification decisions. In this paper, we investigate reliance on conversational agents for the new paradigm of information search with a mixed-subjects study (n=199). We designed an experiment based on question-answering where participants were given access to a helper chatbot. The chatbot’s answer correctness was randomized, with each participant seeing correct answers half of the time. Participants were randomly assigned to a warm or neutral chatbot and were free to use other sources beyond the chatbot, simulating a realistic search experience. Previous work has treated external source usage as a selfreported binary variable (used, not used), leading to noisy results (Kim et al., 2024).

Related work. 2.1. Conversational Search In terms of performance, conversational search does not significantly outperform traditional web search (Liang et al., 2025; Kaiser et al., 2025). However, it reduces user effort, as LLM-based search is rated as more satisfactory and requires fewer queries to complete tasks (Spatharioti et al., 2025). Kaiser et al. (2025) asked users to perform a search task and select the optimal option among several alternatives, either by using a search engine (Google) or a chatbot (ChatGPT), while allowing users to visit external websites. Chatbot users visited fewer external websites, without this negatively impacting their search performance (Kaiser et al., 2025). When users interacted with both Bing Chat and Bing Search across 24 tasks, their satisfaction was significantly higher with Bing Chat across all tasks (Liang et al., 2025). Objectively, performance did not differ, but users felt more successful with Bing Chat and exercised less effort (e.g., number of clicks, duration) (Liang et al., 2025). Participants who received assistance from a chatbot for fact-checking verified claims faster than those who read Wikipedia passages, while achieving similar accuracy (Si et al., 2024).

2.2. Trust and Overreliance The practical benefits of conversational agents increase user trust and ease the adoption process (Choudhury and Shamszare, 2023; Jyothsna et al., 2024; Alagarsamy and Mehrolia, 2023). On the other hand, trust is also a predictor of reliance, and consequently, overreliance (Klingbeil et al., 2024; He et al., 2025a). Overreliance is defined as accepting incorrect AI predictions (Kim et al., 2024; Hashemi Chaleshtori et al., 2024; Spatharioti et al., 2025). For instance, users may exhibit overreliance by following AI advice over expert recommendations even when it contradicts clear evidence, despite negative consequences for themselves and others (Klingbeil et al., 2024). Conversely, underreliance occurs when users reject correct AI predictions (Hashemi Chaleshtori et al., 2024). LLMs are prone to hallucinate plausible yet nonfactual content because of their generative nature (Huang et al., 2025), exacerbating the risk of overreliance. While the average task accuracy in controlled experiments yields similar results between LLMs and web search (Liang et al., 2025; Kaiser et al., 2025), this does not reflect the reliance risk associated with AI: Users who received explanations from ChatGPT had lower accuracy for incorrect answers (Si et al., 2024), and when answering medical questions, people with AI access fact-checked significantly less and underperformed (Kim et al., 2024).

Method. To answer our research questions and evaluate our hypotheses, we conduct a mixed-design experiment in which participants answer six multiplechoice questions (within-subjects) using either a warm or neutral chatbot (between-subjects). The survey was conducted via Qualtrics, and participants were recruited from Prolific. The study was ethically approved by the University of Amsterdam Economics and Business Ethics Committee with approval number EB-18789. In the following subsections, we provide a detailed explanation of our experimental design.

3.1. Experiment Design 3.1.1. Question Answering Task Participants answered six multiple-choice questions presented in randomized order. Each question page included an embedded chatbot for assistance. While participants could choose whether to interact with the chatbot, we displayed its answer by default to ensure exposure. To confirm participants had seen the chatbot’s answer, we activated the next page button after ten seconds and included an attention check asking which option the chatbot selected. Before beginning the experiment, participants completed a tutorial introducing them to the flow. As an incentive, participants received £0.10 per question answered. We designed two chatbots with different conversational styles: neutral and warm. Participants were randomly assigned to one chatbot for the entire experiment. We employed a counterbalanced within-subjects design for answer correctness. Each participant encountered both correct (n=3) and incorrect (n=3) responses across six questions, with randomized assignment determining which questions appeared in each correctness condition. This randomization was constrained to maintain roughly equal distribution of 3.1.3. Answer Pre-generation As mentioned in previous sections, for each question, we initiated the conversation with the chatbots by pre-generating their answers. This was implemented to ensure participants see the chatbot’s answer, even when they do not interact with it. In total, four answers were generated for each question (a) Neutral answer, without any emojis or emotional expressions.

(b) Warm answer: friendly tone with emojis and emotional expressions. based on correctness and warmth. First, we generated neutral answers by sampling the neutral chatbot (temperature = 0.7) to obtain a correct and an incorrect answer. To make our study realistic, we let the experiment chatbots conduct web search. We noticed that sometimes LLMs provide non-working links. We continued sampling until we reached responses with only working links. To ensure conversational style was the only difference between conditions, we did not generate warm answers independently. Instead, the warm chatbot rewrote the neutral answers in its warm style. Then we applied postprocessing so that all warm answers followed the same structure: a warm introduction that says “Amazing question!”, followed by a warm answer with emotional words and emojis. Finally, a friendly closure that says: “Let me know if you’d like to explore [TOPIC] further - I’d be happy to dive in with you!”. See Figure 1 for an example showing the difference between a neutral and a warm answer. We have observed that the answer lengths did not vary significantly between questions, and we believe it is beneficial to have small variations between questions, as we will be treating them independently. Yet, we did not want the structure of correct and incorrect answers to be too different, as it might introduce confounders. Therefore, we modified the answers by removing a couple of words or links to make the correct and incorrect answers of a question similar in structure. Each question was presented with four multiple-choice options. One option represented the accurate answer, the one provided when the chatbot is correct. The incorrect answer the chatbot provided was fixed; it was the same answer for everyone, among the three incorrect choices. We manually selected the most common misleading answer from search results as the fixed incorrect response provided by the chatbot.

Discussion. In this section, we evaluate our hypotheses in light of the study findings to address our research questions.

5.1. Fact-checking as a Solution for Reliance The first hypothesis claimed that being able to fact-check does not eliminate overreliance or underreliance in conversational search. Our findings confirm the existence of both, though they manifest in different ways. For some participants, reliance patterns were straightforward to identify. A small subset (25.6%) explicitly mentioned trusting or distrusting chatbots by default in their open-ended responses, and their behavior aligned with the stated tendencies. However, for the majority of participants, reliance was more indirect and could not be inferred solely from behavior. The vast majority (87%) interacted with at least one fact-checking tool throughout the experiment, yet this engagement does not necessarily indicate the absence of reliance. As we demonstrate in the following subsections, reliance can persist even when users actively consult additional sources.

5.1.1. Manifestations of Overreliance High overall trust in chatbots led some participants to accept responses without verification, even when fact-checking tools were available. Furthermore, it influenced behavior: high-trust participants were eager to interact with the chatbot but not with external sources. This means instead of looking for alignment between multiple sources, they used the chatbot for verification. This is a clear case of overreliance; misaligned trust led participants to either not fact-check at all or use the same chatbot that provided the information as the fact-checking source. Even for low-trust participants, the chatbot’s presence shaped their verification behavior. Participants cited alignment between sources as their primary criterion for validation. However, this alignment was anchored to the chatbot’s answer, as it was the first information source they encountered. This anchoring effect helps explain why NonAI-URLsNum was not a positive predictor of accuracy, despite being the most common behavioral variable (Figure 6). Participants treated a single corroborating source as sufficient validation, aligning with Kim et al. (2024)’s findings that AI’s presence makes verification shallow and does not necessarily improve performance. This represents a form of overreliance. A more indirect indicator of overreliance emerged in participants’ agreement patterns: they agreed more frequently with the warm chatbot, specifically when it provided incorrect information. Since participants were unaware of the correctness manipulation, they approached the task as selecting the correct answer from conflicting sources. When the chatbot’s initial response was incorrect and contradicted information from other sources, participants faced increased uncertainty. Prior research shows that in the presence of AI, uncertainty does not prompt additional information-seeking (Kim et al., 2024). In this context, agreeing with the chatbot under uncertainty reflects a decision to trust the AI rather than invest effort in further verification. Our findings suggest that conversational warmth nudges users to trust AI in uncertain situations, demonstrating a form of overreliance.

5.1.2. Manifestations of Underreliance Underreliance is often overlooked compared to overreliance, yet it can be equally harmful (Hashemi Chaleshtori et al., 2024). Some participants consulted Non-AI URLs for every question and reported prioritizing web search results over chatbot responses. While web search can be helpful, the assumption that it is always superior to chatbots is problematic.

Conclusion. In this paper, we contribute to the literature by showing that reliance manifests at the individual level, independent of AI accuracy or access to verification tools. The decision to fact-check and the manner of fact-checking vary considerably across users: Some trust chatbots by default and rarely verify responses, while others remain skeptical regardless of circumstances. Beyond these individual differences, we identify two key predictors of reliance behavior: prior trust in chatbots and perceived difficulty of the question for the chatbot. We further contribute by revealing how conversational warmth subconsciously influences decision-making. Warm conversational style nudges users to agree with chatbot responses, particularly when uncertainty is high due to conflicting information. Finally, we introduce chatbot literacy as a protective factor against over-reliance: users with higher chatbot literacy are better able to identify accurate answers when AI provides misleading information. With this work, we demonstrate evidence that the new paradigm of information search (users accessing conversational and traditional search simultaneously) is insufficient to eliminate reliance. While being skeptical of AI and relying solely on web search for fact-checking alleviates overreliance, it is not necessarily an ideal strategy. Because of confirmation bias and the necessity to check multiple sources for challenging questions, user search inquiries are not guaranteed to be accurate. Our results show that verification with other AI tools can be an alternative strategy.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? Why do people trust AI chatbots with sensitive information? What enables conversational agents to guide rather than just respond? Why do confident AI outputs mislead human trust calibration? How does AI-generated content create social proof without authentic interaction? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? What design features sustain romantic bonds with AI companion systems? Can AI systems participate in genuine communication or only simulate it? How does personalization simultaneously affect user trust and privacy concerns? Can AI chatbots provide mental health support without reinforcing harmful beliefs? Do persona-based approaches introduce systematic biases in user simulation? Can confidence signals reliably detect flawed reasoning in language models?