An Eye Tracking Study: Are AI Overviews Changing Search Behavior?

Paper · Source
Knowledge After the Web

Source: RMIT University / Microsoft (SIGIR '26) · 2026-07

We conduct a lab eye-tracking study to examine how users interact with search engines that place Generative Artificial Intelligence (GenAI) results above traditionally ranked search results, also known as “ten blue links”, and we use an engagement scale to evaluate their experience. Our aim is to study how users interact with search engine interfaces that incorporate GenAI content, assess users’ willingness to scroll past GenAI content to view the traditional search results, and explore how these interactions differ from existing literature on scanning search engine result pages. We show that GenAI content is changing how people search, but the “golden triangle” remains valid, where the top-left section of the search page attracts the most attention. Searchers are still engaging with the blue links in patterns consistent with the literature; however, they engage significantly more with GenAI content. Finally, we outline future directions to deepen our understanding of search behavior in the era of GenAI.

Introduction. The intersection of Generative Artificial Intelligence (GenAI) and information access has been defined as Generative Information Retrieval (GenIR) [7, 29, 39]. We still lack a clear understanding of how users interact with GenIR systems. Building clear, evidencebased models of user behaviors can improve how we evaluate these systems and their usability. This is why a number of eye-tracking and cognitive user studies were conducted on traditional search engines with the Traditionally Ranked Results also known as “ten blue links” [2, 3, 12, 17, 42]. However, previous findings do not directly apply to modern search engines that incorporate GenIR systems, which often summarize multiple results in the text they generate. Interactive Information Retrieval (IIR) researchers have therefore asked whether GenAI’s diverse interaction possibilities will see widespread adoption or stay a research interest [7]. If so, these new developments in search engine functionality may signal a generational shift in how people search for information. We could then ask, is relying on the traditional “ten blue links” becoming less common, or does it depend on the nature of the tasks at hand? While some studies have explored user interactions with GenIR tools [29, 37, 44], we did not find any lab studies focused on eyetracking or any other physiological devices in this context. By researching human interactions with GenIR systems, we can contribute to developing these systems into true assistants [38, 39]. Search interfaces are constantly evolving, with the design space for search engine interfaces being vast, but interaction techniques and user interfaces with GenIR remain under-researched and poorly understood. Between 2020 and 2024, 750 preprints related to Large Language Models (LLMs) were published on arXiv in the field of Information Retrieval (IR), with only 22 mentioning “user interface” in their abstracts [7]. Information-seeking behavior has changed as search engines have introduced AI-generated summaries and conversational search features in place of simple keyword matching. These features enable more natural interaction and may attract users by saving time, reducing cognitive effort, and improving search efficiency. This shift has contributed to the emergence of GenIR as a distinct research area within IR. Unlike traditional IR systems, which primarily return ranked lists, GenIR systems generate written answers that synthesize information from multiple sources. Many papers in top IR journals and conferences note that user studies in this area remain scarce [7, 16]. To our knowledge, no eye-tracking studies have examined how users interact with these interfaces. This gap is important because eye-tracking is an established method in IR and IIR for studying search behavior [17].

Eye-tracking studies of Search Engine Results Page (SERP) have provided important insights into how users identify relevant or irrelevant documents [20], how their search behavior can be modeled [2, 21, 33], and how attention is shaped by contextual factors [12]. One of the well-known findings in the area is the Golden Triangle Pattern or F-pattern [31], showing that users tend to focus their attention on the top-left portion of the results, even if those results were not relevant. Insights from these user studies informed the development of algorithms and user models [33]. This study examines how users interact with SERP that integrate GenAI content, followed by the Traditionally Ranked Results. It addresses an important gap, as it remains unclear how current search behavior differs from patterns reported before GenAI was incorporated into SERPs. To investigate this, we conduct an exploratory eye-tracking study on SERPs that places GenAI content above the Traditionally Ranked Results.

RQ1: How do users scan and navigate search engine results pages that have Generative AI content placed before the Traditionally Ranked Results? RQ2: To what extent do perceived user engagement levels correlate with users’ gaze behavior on search engine results pages containing Generative AI content before the Traditionally Ranked Results? RQ3: What are the perceived levels of user engagement when interacting with search engine results pages that have Generative AI content placed before the Traditionally Ranked Results?

By addressing our research questions, we contribute the following:

• Demonstrate that users continue to engage with Traditionally Ranked Results in patterns consistent with the literature, but show significantly higher engagement with GenAI content when both appear on the same page. • We release our experimental setup for SERPs with GenAI content and Traditionally Ranked Results, to support reproducibility across other user populations.1

Related work. In this section, we discuss research on Human–GenAI interactions, primarily in IR and non-IR, along with previous eye-tracking studies on how people interact with search engines. In IR, several studies investigate how users interact with GenIR systems [29, 37, 40, 44]. A study [37] uses Bard log submissions from 95 crowd workers, considering factors such as gender, English skills, education level, search skills, and usage of search engines and LLMs. Another User Study [40] methodology consists of user groups from college students and crowd workers from various backgrounds. The study developed a ChatGPT-like interface with GPT-3.5 Turbo and integrated questionnaires and supportive functions, such as Perception Articulation, Prompt Suggestions, and Conversation Explanation, to determine which of those functions supported different user groups. We are also witnessing the publication of various resources and tools in IR to facilitate the study of user behavior in the era of GenAI [28, 46]; however, it’s unclear whether they are suitable tools when running eye tracking studies.

Work from Yang et al. [44] and Liang et al. [29] looked at how users interact with the GenIR system in comparison with the traditional IR system, but the GenAI interaction component was mainly via a chat-based system. A recent study by Wardle et al. [41] researched how people meet their health-related information needs via search engines, including Google’s “AI Overviews” (Powered by Gemini), ChatGPT, and Alexa. The main methodology was structured observation through the think-aloud protocol. It is important to note that the study mentioned that the AI Overviews were rolled back while data were being collected. Xu et al. [43], looked at GenAI content’s influence on public opinion. None of the studies discussed above has utilized eye tracking devices to measure cognitive load. Several eye-tracking studies have extensively analyzed gaze patterns across search engine interfaces [2, 3, 12, 15, 17, 30]. None of these studies explored user interfaces for information access with GenAI either; however, their findings can be compared with ours to examine the generational shift. Abualsaud and Smucker [2] used eye tracking to analyze user behavior while viewing SERP, on both desktop and mobile platforms. Their research tracked participants from the moment they entered a search query to their subsequent actions, which typically involved either clicking a result or refining their query. The study discussed the relationship between user behavior and queries to understand their impact on re-query decisions. Liu et al. [30] utilized eye tracking to analyze fixation patterns and examine user focus on various elements of search interfaces.

Method. To the best of our knowledge, no eye-tracking studies have explored how users interact with search engine interfaces that have GenAI content placed before the Traditionally Ranked Results. In this section, we present the methodology for the user study setup, which will help us answer RQ1–RQ3. We conducted an in-lab user study, collecting gaze data using a Tobii Pro Fusion eye-tracking device and Tobii Pro Lab software. To ensure the validity of our results, we follow the instructions and guidelines provided by Tobii2. We use Qualtrics3 for the surveys, which include the user study tasks 4.

We used a within-subjects design, in which the same participants completed all ten tasks. We collected pre-study participant demographics and characteristics, including age, gender, education, search experience, and prior LLM use. Each task (see below) included a pre-task questionnaire and controlled “backstory” as well as a single central SERP. After each task, participants completed a post-task questionnaire rating their search experience, including perceived relevance of the results and perceived results difficulty. In the exit stage, participants completed the User Engagement Scale (UES) [35], which measures aesthetic appeal, focused attention, perceived usability, and reward. Finally, they answered open-ended questions about their impressions of the SERP, trust in AI overviews Tasks and Searches. Participants completed ten tasks. Each task consisted of three stages (i) completing the pre-task questionnaire, (ii) reading a backstory, and (iii) completing the search task by reviewing a simulated SERP while we logged the interaction behaviors. For each task we used a different search task (i.e., backstory) from the UQV100 test collection5. These backstories (i.e., short scenario descriptions to motivate and contextualize a search task, see Figure 1) represented real-world scenarios in which participants would need to search for information [9]. We selected backstories to provide variety in both topic and word count, and excluded topics related to health, medicine, and politics to minimize potential discomfort. Backstories were categorized by task complexity (see below). Participants were assigned 10 tasks, all participants completed the same set of tasks, and the task order was randomized per participant. For each task, participants first read the backstory and then clicked a link that issued a pre-typed query (see Figure 2 for an example SERP). They reviewed the results on a simulated SERP and could open results pages until they felt their information need was met. Participants were given unlimited time per task and could not switch between tasks once started. The simulated SERP was based on Google’s interface at the time6 and presented GenAI overview or GenAI content, also called “AI Overview”, above the Traditionally Ranked Results. The prewritten queries were selected from an earlier eye-tracking study conducted in early 2024 with different users, during a period when GenAI was becoming more widely adopted. In that study, users wrote queries corresponding to the same set of backstories used in this eye-tracking study. The selection was based on two criteria: their NDCG@10 scores as a measure of their effectiveness [45] and whether they produce AI Overviews from Google. We used user-written queries from the GenAI era to capture more realistic user experiences. The Anserini toolkit was used for retrieval processes7. We used BM25 on the ClueWeb12-B corpus, with parameters set to k1 = 0.9 and b= 0.4 to retrieve 1,000 documents per query. BM25 matched the terms between queries and documents to determine relevance. For this study, we selected NDCG@10 to evaluate participants’ performance and query effectiveness, similarly to [6, 30, 45]. The NDCG@10 values ranged from 0.1 to 0.6, with a range of 0.5 and a standard deviation of 0.1.

The researcher provided a verbal explanation of the study and presented written instructions. Participants reviewed the study protocol and then signed the consent form. To minimize distractions, the researcher left the lab before participants began the tasks.

Discussion. We discuss the implications of our findings, and compare them with previous eye-tracking literature on SERPs. We found that AI Overviews received significantly longer eye fixation and saccade times than Traditionally Ranked Results; however, this was not the case for saccade length. We found that Remember tasks elicited higher fixation and saccade times in AI Overviews, while Understand tasks showed the highest fixation and saccade times in Traditionally Ranked Results and the highest saccade length in AI Overviews. When this was later investigated, it appears that some users may have spent a longer time on Traditionally Ranked Results for Understand tasks in comparison to Remember and Analyze. This may indicate that users prefer the Traditionally Ranked Results rather than AI Overviews for Understand tasks. However, some confounding variables need to be investigated in our future work, such as controlling for complexity by ensuring equal numbers of tasks in each complexity category.

Our aggregated heatmap from our experiments shows the classic “golden triangle” and F-shape of user behavior still holds even though the underlying content our users examined is new compared to past work. Buscher et al. [12] found that when sponsored ads appeared before ten ranked results, the first ranked result received the most attention, not the ads. When less interesting content is returned by a searching system, it would appear that users change their attention behavior. We speculate that our users examined AI Overview content due to interest, not just because the content appeared at the top. Wu et al. [42] examined the effect of presenting a direct answer at the top of a SERP, followed by Traditionally Ranked Results. We computed the proportions of fixation durations on AI Overviews and Traditionally Ranked Results in our eye-tracking study and compared these with proportions reported in prior work, where participants were presented with either only Traditionally Ranked Results or a single direct answer followed by Traditionally Ranked Results (see Figure 6). It is important to note that our study setup, interfaces, and participant sample differed from those in prior work. When one answer was provided, the proportion of fixation duration was 0.50, compared to 0.59 for the AI Overview condition. Participants in the condition with only Traditionally Ranked Results, compared to those in the one-answer condition, spent considerably more time on the first-ranked result (0.31 vs. 0.20). Furthermore, when comparing the one-answer condition with the AI Overview, fixation on the first-ranked result was lower (0.20 vs. 0.09). Overall, we observe a substantial decrease in fixation on the first-ranked result across the three SERP conditions. The proportion of the first Ranked Result (0.09) in the AI Overview Condition is comparable to the proportions of the 4th and 5th Ranked Result positions (0.11 and 0.07) in the condition with only Traditionally Ranked Results. This may suggest that the 1st Ranked Result is effectively shifted to the 4th or 5th position when AI Overviews are placed above the Traditionally Ranked Results. The F-pattern remains consistent, with the top-left position continuing to attract the most attention, as observed in the literature, but that position is occupied by an AI Overview instead of the first-ranked result, as further supported by the aggregated fixation heatmaps in our findings. First-ranked results attracted the most attention and clicks, with the fastest arrival time, while attention and clicks decreased and arrival time increased down the ranked list. This is consistent with

Conclusion. This eye tracking study examines how users interact with a SERP that presents GenAI content at the top of the page, immediately above the Traditionally Ranked Results. The study provides empirical evidence that AI Overviews are changing the way people search, but not to the extent that the golden triangle or the F-pattern observed in SERP are no longer valid. Users still engage with the Traditionally Ranked Results in patterns consistent with the literature [13, 17]; however, they engage more with AI Overviews than with Traditionally Ranked Results. Despite our study providing evidence that users still interact with the Traditionally Ranked Results in patterns similar to those reported in the literature, further research is required to understand the extent to which those patterns are similar. As for the qualitative analysis of trustworthiness and utility, we observed mixed results, with neither AI Overviews nor Traditionally Ranked Results perceived as more trustworthy or useful than the other. Given the rapid evolution of interfaces integrating GenAI, future studies with our setup could explore AI Overviews that contain multiple links as sources and that appear in different SERP positions. Future work could investigate GenIR systems across specific tasks, diverse user groups, or different ecologically valid settings. It could also group users by survey responses, UES ratings, and qualitative feedback, such as reported trust in the system and perceived usefulness. This would clarify how search behavior and gaze patterns vary across groups and which factors most strongly shape these patterns. This work would help the IR community in making recommendations for optimizing GenAI systems based on user preferences and behaviors, developing a framework of how people are influenced to search in the era of GenAI, and working toward making these systems true assistants.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Why do confident AI outputs mislead human trust calibration? How should humans and AI agents share control and decision-making? When do simpler collaborative filtering approaches outperform complex LLM recommenders? How should retrieval strategies adapt to multi-step reasoning demands? When should retrieval systems decide to fetch new information? How should recommendation systems balance individual preference and diversity? Why do abstract preferences outperform episodic memories in personalization?