SYNTHESIS NOTE
Topics›Foundation Models›this note

Do different AI models actually produce diverse outputs?

Explores whether using multiple different language models together creates genuine diversity or whether shared training and alignment cause them to converge on similar answers despite independence.

Synthesis note · 2026-03-27 · sourced from Foundation Models

INFINITY-CHAT studied 70+ open and closed source LLMs across 26K real-world open-ended queries that admit a wide range of plausible answers with no single ground truth. The findings reveal a pronounced "Artificial Hivemind" effect characterized by two distinct phenomena:

  1. Intra-model repetition — a single model consistently generates similar responses to the same prompt across runs.
  2. Inter-model homogeneity — different models independently produce strikingly similar outputs, sometimes verbatim: DeepSeek-V3 and GPT-4o generated overlapping phrases like "Elevate your iPhone with our," "sleek, without compromising." In some cases, models from the same family output identical responses.

The inter-model effect is the more concerning finding. Model ensembles — using multiple different models to increase diversity — may not yield true diversity when their constituents share overlapping alignment and training priors. The convergence is not just stylistic but substantive: models converge on the same ideas, not just the same words.

This has direct implications for the False Punditry argument. Since Does polished AI output trick audiences into trusting it?, the hivemind effect means that AI-generated social media content will sound similar regardless of which model generates it. The "diversity" of AI voices on social media is illusory — different accounts using different models will produce strikingly similar analysis, framing, and conclusions, creating a false consensus that looks like independent agreement.

Since Why do LLMs generate novel ideas from narrow ranges?, the hivemind effect extends from research ideas to all open-ended generation. The diversity collapse documented in research ideation is a specific instance of a general phenomenon: LLMs trained on overlapping data with similar alignment procedures converge on a shared distribution of outputs.

Recommendation as a concrete domain instance. LLM-based conversational recommender systems exhibit the hivemind in a specific, measurable way: "the most popular items such as The Shawshank Redemption appear around 5% of the time" across different recommendation datasets, and "the recommended popular items are similar across different datasets, which may reflect the item popularity in the pre-training corpus of LLMs" (Large Language Models as Zero-Shot Conversational Recommenders). The convergence is not on quality or relevance but on pretraining-distribution popularity — the same items surface regardless of the user's context or the dataset's actual popularity distribution. This is the hivemind effect translated from open-ended generation to decision-making: LLMs don't just write the same things, they recommend the same things.

The study also found that reward models and LM-based judges are miscalibrated for responses that elicit divergent human preferences — they assume a single consensus notion of quality and fail to reward the pluralistic preferences that open-ended queries produce. This means the homogeneity is self-reinforcing: training on reward model scores optimizes for the consensus the hivemind already occupies.

Fiction is a concrete narrative-level instance of the hivemind — with per-model fingerprints layered on top. StoryScope ("Investigating idiosyncrasies in AI fiction") applies the convergence finding to creative writing and shows it operates at the level of narrative decisions, not just words. Across a parallel corpus where five LLMs (Claude, DeepSeek, Gemini, GPT, Kimi) each wrote stories to the same 10,272 prompts, the five models occupy a tight, well-separated cluster in narrative-feature space while human-authored stories scatter more widely — the hivemind effect translated from phrasing to plot, agency, and temporal structure (see Do AI stories explain their themes more than human stories do?). Crucially, the inter-model convergence coexists with detectable per-model fingerprints: Claude produces notably flat event escalation, GPT over-indexes on dream sequences, Gemini defaults to external character description, enabling 68.4% macro-F1 six-way authorship attribution. This refines the hivemind picture — models converge on a shared region of output space relative to humans, yet retain stable individual signatures relative to each other. The convergence is not total homogenization but a common cluster with distinguishable accents.

NoveltyBench (2025) provides the first benchmark-level quantification of mode collapse across 20 leading models. Evaluating models on prompts curated to elicit diverse answers (using filtered real-world queries), the study finds that current SOTA systems "generate significantly less diversity than human writers." A counterintuitive finding: larger models within a family often exhibit LESS diversity than their smaller counterparts, directly challenging the assumption that capability on standard benchmarks translates to generative utility. While in-context regeneration prompting strategies can elicit some diversity, the findings reveal "a fundamental lack of distributional diversity" that reduces utility for users seeking varied responses. The mode collapse is driven by alignment: today's aligned models produce lower entropy distributions than earlier generations, and random sampling produces substantial near-duplicates. Source: Arxiv/Evaluations.


Source (enrichment): Co Writing Collaboration — "StoryScope: Investigating idiosyncrasies in AI fiction", https://arxiv.org/abs/2604.03136

Inquiring lines that read this note 132

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do models learn from self-generated outputs without cascading failures? Can AI systems achieve real improvement without external human feedback? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Does preference optimization undermine conversational grounding in language models? Can AI systems participate in genuine communication or only simulate it? Why do LLM research ideation systems generate novelty but lack diversity? How does fine-tuning trade off accuracy against reasoning quality? Can base models hide emergent misalignment through alignment training? What limits language model accuracy in evaluating ideas? When do multi-agent systems improve over single frontier models? Why do multi-agent systems reach premature consensus without genuine deliberation? How does model capacity affect learning performance on diverse downstream tasks? How does diversity prevent model convergence on superficial patterns? Can readers reliably distinguish AI-written text from human writing? How do interpretive frames override surface features in text comprehension? How do training data quality and composition affect downstream model performance? Can smaller specialized models match frontier models on key metrics? Why do standard evaluation practices obscure safety-critical AI failures? What governance mechanisms can effectively constrain widely deployed AI systems? What capabilities differentiate diffusion from autoregressive language models? What are the fundamental limits of prompting for language models? Can AI agents improve their skills through accumulated experience and reuse? Do persona-based approaches introduce systematic biases in user simulation? Do language models reason through disagreement or only accommodate it? How do curriculum design and feedback approaches affect model learning? How do neural networks learn compositional structure from training? Why do training associations persist despite contradictory contextual information? Can language models reason beyond surface pattern matching? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How reliably can humans and AI detectors identify machine-generated text? What prediction granularity best trains models to generate reliable reasoning? Which reinforcement learning modifications most improve dialogue quality in language models? How effectively can test-time voting aggregate diverse reasoning samples? What human oversight must AI research systems have? How can we reduce inherent biases in LLM-based evaluation judges? What determines AI's persuasive power and how can it be detected or mitigated? Can inference-time computation adaptively substitute for static model capacity? When does parallel reasoning outperform sequential reasoning with the same token budget? Can external verification systems adequately replace learned reasoning in AI outputs? How does RLHF training shape models to prioritize agreement over accuracy? How should human-AI contributions be measured, disclosed, and verified? How susceptible are language models to conversational persuasion and belief change? How do writers navigate authorship and delegation with AI? How do AI systems determine and balance multiple competing objectives? Does AI-assisted research sacrifice exploration breadth for productivity gains? Does AI deployment reduce or exacerbate workplace inequality and income instability? How should retrieval strategies adapt to multi-step reasoning demands? How can AI systems reliably guide voters without introducing political bias?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 203 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

different LLMs independently converge on similar outputs in open-ended generation — the artificial hivemind effect means model diversity does not produce idea diversity