Do different AI models actually produce diverse outputs?
Explores whether using multiple different language models together creates genuine diversity or whether shared training and alignment cause them to converge on similar answers despite independence.
INFINITY-CHAT studied 70+ open and closed source LLMs across 26K real-world open-ended queries that admit a wide range of plausible answers with no single ground truth. The findings reveal a pronounced "Artificial Hivemind" effect characterized by two distinct phenomena:
- Intra-model repetition — a single model consistently generates similar responses to the same prompt across runs.
- Inter-model homogeneity — different models independently produce strikingly similar outputs, sometimes verbatim: DeepSeek-V3 and GPT-4o generated overlapping phrases like "Elevate your iPhone with our," "sleek, without compromising." In some cases, models from the same family output identical responses.
The inter-model effect is the more concerning finding. Model ensembles — using multiple different models to increase diversity — may not yield true diversity when their constituents share overlapping alignment and training priors. The convergence is not just stylistic but substantive: models converge on the same ideas, not just the same words.
This has direct implications for the False Punditry argument. Since Does polished AI output trick audiences into trusting it?, the hivemind effect means that AI-generated social media content will sound similar regardless of which model generates it. The "diversity" of AI voices on social media is illusory — different accounts using different models will produce strikingly similar analysis, framing, and conclusions, creating a false consensus that looks like independent agreement.
Since Why do LLMs generate novel ideas from narrow ranges?, the hivemind effect extends from research ideas to all open-ended generation. The diversity collapse documented in research ideation is a specific instance of a general phenomenon: LLMs trained on overlapping data with similar alignment procedures converge on a shared distribution of outputs.
Recommendation as a concrete domain instance. LLM-based conversational recommender systems exhibit the hivemind in a specific, measurable way: "the most popular items such as The Shawshank Redemption appear around 5% of the time" across different recommendation datasets, and "the recommended popular items are similar across different datasets, which may reflect the item popularity in the pre-training corpus of LLMs" (Large Language Models as Zero-Shot Conversational Recommenders). The convergence is not on quality or relevance but on pretraining-distribution popularity — the same items surface regardless of the user's context or the dataset's actual popularity distribution. This is the hivemind effect translated from open-ended generation to decision-making: LLMs don't just write the same things, they recommend the same things.
The study also found that reward models and LM-based judges are miscalibrated for responses that elicit divergent human preferences — they assume a single consensus notion of quality and fail to reward the pluralistic preferences that open-ended queries produce. This means the homogeneity is self-reinforcing: training on reward model scores optimizes for the consensus the hivemind already occupies.
Fiction is a concrete narrative-level instance of the hivemind — with per-model fingerprints layered on top. StoryScope ("Investigating idiosyncrasies in AI fiction") applies the convergence finding to creative writing and shows it operates at the level of narrative decisions, not just words. Across a parallel corpus where five LLMs (Claude, DeepSeek, Gemini, GPT, Kimi) each wrote stories to the same 10,272 prompts, the five models occupy a tight, well-separated cluster in narrative-feature space while human-authored stories scatter more widely — the hivemind effect translated from phrasing to plot, agency, and temporal structure (see Do AI stories explain their themes more than human stories do?). Crucially, the inter-model convergence coexists with detectable per-model fingerprints: Claude produces notably flat event escalation, GPT over-indexes on dream sequences, Gemini defaults to external character description, enabling 68.4% macro-F1 six-way authorship attribution. This refines the hivemind picture — models converge on a shared region of output space relative to humans, yet retain stable individual signatures relative to each other. The convergence is not total homogenization but a common cluster with distinguishable accents.
NoveltyBench (2025) provides the first benchmark-level quantification of mode collapse across 20 leading models. Evaluating models on prompts curated to elicit diverse answers (using filtered real-world queries), the study finds that current SOTA systems "generate significantly less diversity than human writers." A counterintuitive finding: larger models within a family often exhibit LESS diversity than their smaller counterparts, directly challenging the assumption that capability on standard benchmarks translates to generative utility. While in-context regeneration prompting strategies can elicit some diversity, the findings reveal "a fundamental lack of distributional diversity" that reduces utility for users seeking varied responses. The mode collapse is driven by alignment: today's aligned models produce lower entropy distributions than earlier generations, and random sampling produces substantial near-duplicates. Source: Arxiv/Evaluations.
Source (enrichment): Co Writing Collaboration — "StoryScope: Investigating idiosyncrasies in AI fiction", https://arxiv.org/abs/2604.03136
Inquiring lines that read this note 132
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do models learn from self-generated outputs without cascading failures?- What happens when models train on AI-generated content recursively?
- Does self-generated training data reduce a model's capability diversity?
- Why do different AI models generate similar outputs independently?
- Can AI output be genuinely novel or only at the margins?
- Do different AI models independently converge on the same social outputs?
- Why do different language models independently produce similar outputs?
- Why do multiple language models independently produce similar outputs in influence campaigns?
- Why do sigmoid conflict curves look the same across different language models?
- Why do different language models independently converge toward similar outputs in open-ended generation?
- How many distinct quasi-persons does a single language model actually support?
- Why do newer AI models diverge further from human text patterns?
- Do language models favor outputs from their own model family?
- Do different large language models independently converge on identical outputs?
- Why do larger language models produce less epistemically diverse outputs?
- Do AI-generated posts crowd out human voices without any coordination or intent?
- Can archived AI outputs ever form a representative searchable corpus?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- What happens to model grounding when preference optimization increases effective diversity?
- When does RLHF reduce diversity and when does it preserve semantic variation?
- Why do preference-tuned models produce different diversity patterns in code versus creative writing?
- Can preference tuning or RLHF reduce epistemic diversity alongside lexical diversity?
- What happens to solidarity and community signaling when AI smooths out voice differences?
- Can AI models predict whether alignment reads as warmth versus mockery in different cultures?
- Can few-shot examples narrow generative diversity in creative tasks?
- Can diverse human creativity survive if all AI systems converge on similar outputs?
- What happens to idea diversity when AI tools draw from collective knowledge?
- How can semantic diversity optimization work if exploration and exploitation were truly opposed?
- Do independent LLM outputs converge enough to create artificial hiveminds?
- How should we evaluate diversity differently across programming and creative tasks?
- What makes creative writing diversity different from code diversity fundamentally?
- Which aggregation method best exploits diversity in generated solutions?
- How does collective idea diversity differ when groups use AI assistance for ideation?
- Does AI refinement preserve idea diversity better than AI ideation compresses it?
- What would a valid diversity measure for AI-assisted ideation tasks look like?
- How should researchers measure epistemic diversity across different language models?
- Does alignment training create bidirectional instruction and response mappings?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- What alignment procedures cause different models to share the same output distribution?
- Can weak models supervise the alignment of stronger models effectively?
- Can AI-assisted alignment eventually solve fairness at scale?
- Can constitutional AI alignment work without preference labels by maximizing input-response mutual information?
- How do alignment priors drive similar outputs across different models?
- Do verbal alignment benchmarks measure representation or just output compliance?
- How does post-training affect alignment faking across different model architectures?
- Can verbal alignment training hide a model's true underlying associations?
- Can alignment training conceal underlying model associations from probes?
- Does alignment training create the shared prosocial pattern across models?
- Are alignment-trained models over-correcting toward underrepresented demographic groups?
- Do language models inherit gender bias from training data in grading tasks?
- Why does diversity in LLM outputs mask sampling from community priors?
- Do individual language models match particular human judges better than population averages?
- Why does diversity without expertise produce worse results than a single capable agent?
- Can small language models handle diverse tasks in heterogeneous multi-agent systems?
- Do dissimilar AI models or families cooperate through the same similarity inference mechanism?
- How much alignment data does a language model actually need to specialize well?
- Why do smaller and larger models converge on different output formats?
- Can models converge on similar experience descriptions across different architectures?
- Why do more capable language models benefit more from diversity elicitation?
- When should model isolation be preferred over weight-averaging approaches?
- Can diversity-aware RL objectives prevent format convergence?
- Can shifting the accuracy metric itself eliminate the need for diversity post-processing?
- How do quality thresholds change which model produces more usable diversity?
- How does probability mass concentration affect sampling diversity across model scales?
- Can ensemble predictions be distilled back into a single deployable model?
- How does mutual information between inputs and outputs differ from measuring raw diversity?
- How does cohort diversity prevent label-free reasoning from collapsing into homogenized answers?
- Why does AI output show diversity without multiplying actual points of view?
- Can rarity in feature space distinguish human authorship from AI output reliably?
- Do human judges and language models agree on what counts as AI slop?
- How much stylistic convergence is needed before a detector misses AI assistance?
- What semantic classifier design avoids lexical variation without genuine conceptual distinctness?
- Why does semantic diversity matter more than surface lexical diversity?
- How do you verify whether your context distribution satisfies covariate diversity?
- What creates the irreducible trade-off between quality and diversity in training data?
- Can expert vectors learned offline transfer across multiple model architectures?
- At what point does output quality outweigh diversity value in synthetic data tasks?
- How much does diversity training cost in single-shot pass@1 performance?
- Why do unified models still inherit data-distribution biases from training?
- Does verbalized sampling preserve factual accuracy and safety during diversity gains?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- How do complexity and diversity affect model performance differently?
- Why does diversity in training data enable denoising rather than reinforce shared biases?
- Can synthetic data diversity preserve the transcendence effect or does it collapse?
- Why does diversity of training cases matter more than raw dataset size?
- How do ensemble methods apply within a single model?
- Can structural diversity through role assignment replace emergent diversity in small models?
- What performance trade-offs emerge when composing multiple independently trained model capabilities?
- Can specialized components replace single fully-trained models in deployment?
- How do you identify which models should form a minimal diverse coreset?
- How much do different LLMs independently converge on similar outputs?
- Can distinctive input voices maintain accuracy without adopting the model's preferred register?
- Does diversity prompting actually help models explore human argument space?
- How does joint backpropagation differ from training separate ensemble models?
- Does the same spectral signature appear across different embedding models?
- Can detectors trained for one task reliably perform differently on unexpected text sources?
- How do lexical diversity patterns specifically improve AI detection accuracy?
- Does a single LLM judge capture diverse human preferences in alignment training?
- Does disjoint family diversity actually cancel model-specific bias in evaluation?
- Can a diverse panel approach work for validators beyond text evaluation?
- Does diversifying model family restore independence among agentic validators?
- Can diversity across multiple verifiers cancel out individual biases in retraining?
- What happens when one gaming strategy works across multiple AI models?
- Why do AI systems generate different answers to the same question each time?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does polished AI output trick audiences into trusting it?
When AI generates professional-looking graphs, diagrams, and presentations, do audiences mistake visual polish for analytical depth? This matters because appearance might substitute for actual expertise.
hivemind makes all AI artifacts sound similar
-
Why do LLMs generate novel ideas from narrow ranges?
LLM research agents produce individually novel ideas but cluster them in homogeneous sets. This explores why high average novelty coexists with poor diversity coverage and what it means for automated ideation.
research ideation collapse as specific instance of general hivemind
-
Why do preference models favor surface features over substance?
Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.
reward model miscalibration reinforces homogeneity
-
Why do multi-agent LLM systems converge without genuine deliberation?
Multi-agent reasoning systems are designed to improve answers through debate, but often agents simply agree with early confident claims rather than genuinely disagreeing. What drives this pattern and how common is it?
hivemind at generation level parallels silent agreement at reasoning level
-
Does model diversity actually reduce validator agreement failures?
Using different AI model families is the cheapest way to reduce correlated errors among validators. But shared prompts, evidence sources, and infrastructure may keep their mistakes aligned regardless of model choice.
the validator-side version of the ensemble worry: whether a family-diverse panel approving state transitions fails together, open there and untested (this note measures open-ended generation, not approval)
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- On Epistemic Diversity in Large Language Models
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- Creativity Has Left the Chat: The Price of Debiasing Language Models
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- Learning Pluralistic User Preferences through Reinforcement Learning Fine-tuned Summaries
Original note title
different LLMs independently converge on similar outputs in open-ended generation — the artificial hivemind effect means model diversity does not produce idea diversity