Experimental evidence of the effects of large language models versus web search on depth of learning

Paper · Source
Knowledge After the Web

Source: PNAS Nexus (Melumad, Yun) · 2025-10-28

Since their public release in 2022, large language models (LLMs) such as ChatGPT have become an increasingly pervasive tool for acquiring information (1–4). A defining feature of LLMs is the format in which their results are displayed: whereas traditional online search presents a series of web links that users must navigate and distill on their own, LLMs complete this process on behalf of users, providing automatic syntheses of vast amounts of information (5, 6). Hence, LLMs enable users to learn about topics faster and with less effort than in traditional web search (6, 7). It is perhaps unsurprising, then, that major search engines such as Google now offer “AI Overviews” of their standard search results, putting LLMs at the forefront of search while making web links a secondary resource (5).

While LLMs afford obvious efficiency gains for users, we propose that this greater ease may come at a cost: that of reducing the depth of knowledge, and originality of thought, gleaned from one's search in certain contexts. Specifically, while the need to navigate and summarize different information sources might make learning through web search more effortful, it can offer the often-overlooked benefit of constructing deeper, and more unique, knowledge structures (8, 9)—something that is less likely when various information sources are already synthesized for the user by an LLM. As a result, individuals who learn about a topic from LLM syntheses (vs. standard web links) may at times emerge feeling that they have learned less about it, and, in turn, generate ideas about it that are shallower—for example, by forming advice for others on the subject that is sparser, more generic, and, ultimately, less informative and persuasive to others.

We tested these ideas across a series of experiments in which participants were randomly assigned to learn about a topic either from standard Google web links or LLM syntheses (e.g. ChatGPT) and were then asked to create advice on the subject based on what they learned. Results from seven online and laboratory experiments (n = 10,426) yield robust evidence that in this task context, participants reported developing shallower knowledge when learning about a topic from LLM summaries (vs. web links), and that this occurred because those who used an LLM exerted less effort in learning from its synthesized responses compared with those who gathered and distilled information themselves through web links. As a result, when they subsequently wrote advice on the subject, those who learned through an LLM (vs. web search) were not only less invested in forming their advice, but they also created content that was objectively sparser and more generic—such that recipients found it to be less informative and were less likely to adopt it. These results were robust across different search topics and LLM tools and held even when, for example, LLM summaries were augmented by real-time web links.

A central thesis of this work is that whereas LLMs provide a faster route to finding answers than web search, by doing so they inhibit a process that can be instrumental to learning: the self-guided exploration of different information that requires original synthesis. The value of self-directed knowledge acquisition for skill development has been explored in the literature on “search-as-learning,” which argues that when we engage in traditional web search on platforms like Google, we do more than accumulate facts—we also develop knowledge structures through an iterative process of posing queries, gathering and interpreting information from different websites, and then assembling this knowledge into a cohesive whole (8, 9, 25, 26). Thus, the process of “sensemaking” through web search can be highly dynamic for users, marked by the recursive process of synthesizing and revising (27, 28). In contrast, we argue that since LLMs are designed to perform such sensemaking on behalf of the user, this critical ingredient in learning is often diminished relative to gathering information from web links.

In contrast to web search, when learning from LLM summaries users no longer need to exert the effort of gathering and distilling different informational sources on their own—the LLM does much of this for them. We predict that this lower effort in assembling knowledge from LLM syntheses (vs. web links) risks suppressing the depth of knowledge that users gain, which subsequently affects the nature of the advice they form on the topic for others. Specifically, after learning about a subject via LLM summaries (vs. web links), users will be less invested in writing the advice and, more importantly, their advice will reflect shallower knowledge, such as by being sparser and more generic (34, 35). As a result, recipients of this advice will find the advice to be less informative and will ultimately be less willing to adopt it. We illustrate these ideas in the conceptual model in Fig. 1.

We report the results of seven total experiments that test these predictions, including four experiments in the main text and three in the Supplementary material. The first preregistered experiment provides an initial test of our predictions in a naturalistic setting where participants are asked to learn about a particular topic by engaging interactively with either Google or ChatGPT, and then to provide advice to a friend based on what they learned from their search. In the second preregistered experiment, we report a more conservative test of our predictions by exposing participants to search results containing the same set of facts, varying only whether it is presented in the format of an LLM summary or a set of linked websites. In the third experiment, we test for the robustness of the effects in a laboratory setting, this time by holding constant the search engine—Google search—and varying whether participants learned about a topic through standard Google search results or Google's LLM synthesis (“AI Overview”) presented at the top of the standard search results page. Finally, in the fourth experiment, we explore the downstream consequences of these effects by presenting advice written by participants in a prior experiment to an independent set of “recipients”—blind to the original search platform used to learn about the topic—and examining their willingness to adopt the advice. In the Supplementary Material, we report the results of three preregistered replications that, for example, test for the robustness of the basic effects to the inclusion of real-time web links in LLM syntheses and to a search topic that was of high personal relevance.

Participants were 1,136 membersa of the Prolific panel who were randomly assigned to learn about a topic using either ChatGPT or Google and then to write advice for their friend on it (Mage = 42.40, SD = 13.46; 46% female, 53% male, 1% nonbinary). For the first experiment, we built an in-house platform that allowed participants to conduct actual ChatGPT or Google searches while enabling us to record their submitted queries. Per the preregistered criteria, we excluded 32 participants for failing attention checks, resulting in a final sample of 1,104 participants (Mage = 42.40, 53% female). The preregistration for the study is available at https://aspredicted.org/D1G_BRC.

First, the results support the predicted effects of the search platform on the amount of effort invested in learning from search, as well as on the depth of knowledge participants reported acquiring from that search. As expected, participants who used ChatGPT spent less time on the search task than those using Google search [seconds: MGoogle = 742.81, MGPT = 585.41; F(1, 1,102) = 44.61, P < 0.001], suggesting that learning from LLM syntheses involved less effort than learning from standard web search results. It is worth noting that participants in the ChatGPT condition submitted a similar number of queries on average as those in the Google condition—with 2.06 prompts submitted to ChatGPT and 2.15 queries submitted to Google on average [F(1, 870) = 0.71, P = 0.401]—implying that the lower amount of time participants spent during the ChatGPT (vs. Google) search task was not driven by lower interactivity with ChatGPT compared with Google, but rather by less engagement with the search results.

Importantly, LLM use also suppressed participants' reported depth of learning on the topic compared with traditional web search: those who used ChatGPT (vs. Google) to learn about the topic at hand (how to plant a vegetable garden) reported that they learned fewer new things about the subject [MGoogle = 3.86, MGPT = 3.43; F(1, 1,102) = 36.04, P < 0.001], felt a lower sense of personal ownership over the knowledge they gained from their search [MGoogle = 3.55, MGPT = 3.36; F(1, 1,102) = 6.82, P = 0.009], and thought their search yielded less comprehensive information about the topic [MGoogle = 4.27, MGPT = 4.02; F(1, 1,102) = 20.84, P < 0.001].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Does AI assistance erode cognitive skills while inflating perceived competence? Why does polished AI output gain credibility despite fundamental verifiability problems? How can we detect and account for LLM involvement in academic writing? How do interpretive frames override surface features in text comprehension? What prevents LLMs from applying their reasoning knowledge to improve outputs? How can AI systems reliably guide voters without introducing political bias? How should retrieval strategies adapt to multi-step reasoning demands? Can AI chatbots provide mental health support without reinforcing harmful beliefs? How does AI adoption reshape collaboration patterns in knowledge work? When do simpler collaborative filtering approaches outperform complex LLM recommenders?