SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Are large language models becoming more epistemically diverse?

This research asks whether LLMs are converging toward narrow answer spaces or developing broader epistemic range over time. The question matters because it challenges the assumption that LLM homogeneity is inevitable or unchanging.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

The paper runs "the first systematic study of epistemic diversity in LLMs across time and cultural context," testing 27 LLMs on 155 topics spanning 12 countries, producing 1.7M responses decomposed into 70M individual claims. Against "the dominant paradigm... that overall LLM diversity is low," measured only "with respect to a single point in time," the longitudinal result is that "epistemic diversity has increased substantially over the past three years, a positive counter to recent diversity pessimism." But the paper immediately qualifies its own good news: "despite progress, every system in our study is less diverse than a search baseline," and that gap "is not uniform." Retrieval-augmented generation (RAG) raises diversity; model scale works the other way; and for country-specific topics, "LLM parametric knowledge systematically reflects English over local-language knowledge."

The method explains how a trend claim becomes measurable at all. Responses are decomposed into individual claims, claims are clustered into semantically equivalent "meaning classes," and diversity is scored with Hill-Shannon diversity, an ecology metric normally used for species counts. This lets the authors compare diversity across models, time, and topic rather than eyeballing homogeneity in a single snapshot. The unevenness has a logic in the paper's own terms: RAG helps because it grounds a response in retrieved sources, but "RAG has an uneven effect" since diverse retrieval sources exist unevenly by language — the paper notes the USA "see[s] more benefit due to a greater diversity in their RAG sources, which are not available for many languages." Larger models being "counterintuitively" less diverse is reported as consistent with contemporary work (cited as Zhang et al. 2025 and Rassin et al. 2024) but not causally explained here. The English-over-local-language gap follows from where parametric knowledge comes from: training corpora skew English, so a model's internal knowledge of a country-specific topic draws more from English sources than from that country's own language, even against local-language Wikipedia as a comparison point.

This complicates the library's working picture of homogenization. Do frontier LLMs actually explore the full space of valid answers? measures the same kind of narrowness — using the same Hill-Shannon lineage — but only at one point in time, in two closed domains (professions, proofs); this paper adds the missing time axis and an external search baseline, and finds the trend line is improving rather than flat. Do different AI models actually produce diverse outputs? is likewise a snapshot finding of convergence; read alongside this paper, hivemind-style convergence may describe a moment that is already easing rather than a fixed property of LLMs. And against Do large language models narrow human expression and thought?, which argues narrowing from training statistics and shared reliance, this paper's direct measurement is a check on that argument: narrowing is real relative to search, but it is not static, and RAG and scale push it in opposite directions.

The excerpt gives the shape of the findings without the numbers that would let a reader judge their size: no effect sizes for the three-year increase, no list of which countries gain most or least from RAG beyond the USA example, and no mechanism for why larger models score lower beyond a citation to other papers reporting the same pattern. It also does not say whether "increased diversity" in LLM outputs maps onto any check of factual accuracy — diversity and correctness are separate axes, and this measure is silent on the second. What the data support, at the strength given, is narrower than "LLMs are catching up to search": a real and measured easing of homogenization, still short of search, and distributed unevenly enough that the countries and languages already underserved by the web are the same ones underserved by the models trained on it.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Why do LLM research ideation systems generate novelty but lack diversity?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 94 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM epistemic diversity has grown over three years yet still trails a search baseline — unevenly across RAG, model size, and language