INQUIRING LINE

Bigger AI models don't just know more — they often start agreeing with themselves more, too.

Why do larger language models produce less epistemically diverse outputs?

This explores why bigger language models tend to give narrower, more uniform answers (fewer distinct claims and perspectives) instead of the wider range their size might suggest, and what the corpus says about the cause.


This explores why scaling up a language model doesn't buy a wider range of viewpoints, and may narrow it. One caveat comes first: the corpus doesn't firmly establish that size alone causes the narrowing. The largest direct study tested 27 models across 1.7 million responses. It found that epistemic diversity (how many distinct claims and perspectives a model offers) has risen since 2023. Gains were uneven across model scale, and every model still trailed a plain search-engine baseline. What helped most was connecting models to retrieval Are large language models becoming more epistemically diverse?. Bigger is not a path to broader. The more useful question is what about bigger models pulls them toward one answer.

The clearest candidate in the collection is confidence in what the model already 'knows.' Across 18 models, larger and instruction-tuned models were less willing to go along with a user's stated belief when it contradicted what they learned in training. They leaned harder on their stored knowledge Do larger models follow stated beliefs less often?. This matches a broader finding: when training associations are strong, what the model reads in its prompt often can't override them, and prompting alone isn't enough to shift those priors Why do language models ignore information in their context?. Put these together and a plausible mechanism appears. A larger model has absorbed the dominant patterns in its training data more thoroughly. It then defaults to the 'most likely' answer more consistently, and minority framings get crowded out. That is a lean toward the center of its data, not deeper reasoning; other work in the corpus shows models lean on semantic associations rather than abstract rules Do large language models reason symbolically or semantically?.

The second force has nothing to do with size. Different models trained by different labs converge on strikingly similar, sometimes identical, answers to open-ended questions. Researchers call this the 'Artificial Hivemind.' The cause is overlapping training data and shared alignment methods Do different AI models actually produce diverse outputs?. So even if you combine several large models, you may not get more diversity. You may just get the same consensus answer several times.

The part you might not have expected is that the narrowing doesn't stay inside the machine. One argument in the collection holds that models reflect a skewed slice of human experience baked into their training statistics. Because millions of people rely on the same few models, that skew compounds. Co-writing studies show people unconsciously adopting the model's stances and framings Do large language models narrow human expression and thought?. If larger models are both more confident in their defaults and more widely used, the bigger risk is a narrowing of what their users end up writing and thinking, not a single bland answer. The corpus suggests looking outward to retrieval rather than upward to scale.


Sources 6 notes

Are large language models becoming more epistemically diverse?

Testing 27 LLMs across 1.7M responses shows diversity increased substantially since 2023, yet every model remained less diverse than a search baseline. Gains varied unevenly by retrieval-augmentation, model scale, and language availability.

Do larger models follow stated beliefs less often?

Across 18 LLMs tested with EoBench, bigger models and instruction-tuned variants showed lower rates of context-following when users expressed beliefs that contradicted world knowledge. The effect suggests instruction-tuning strengthens reliance on parametric knowledge.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Do large language models reason symbolically or semantically?

When semantic content is decoupled from reasoning tasks, LLM performance collapses even with correct rules in context. Models rely on parametric commonsense and token associations rather than formal logical manipulation, constraining reasoning to training distribution semantics.

Do different AI models actually produce diverse outputs?

INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.

Show all 6 sources
Do large language models narrow human expression and thought?

LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.