Judged one idea at a time, AI can look more novel than experts, but does its output stay varied or keep converging?
Can AI systems generate diverse hypotheses or do they collapse toward similar ideas?
This explores whether AI can actually produce a wide spread of ideas and hypotheses, or whether its output drifts toward the same few answers, whether you're asking one model, many models, or a team of AI agents.
This explores whether AI produces a real spread of ideas or keeps landing on the same few. The corpus gives a surprising answer: both are true, at different levels. Judged one idea at a time, AI output can look more novel than what experts produce. In a blinded study with 100+ NLP researchers, LLM-generated research ideas were rated more novel than the experts' own ideas, though slightly less feasible Do language models generate more novel research ideas than experts?. The likely reason is that experts know what tends to fail and steer away from it, while a model combines concepts without that caution. AI can also land on the *right* idea. One AI co-scientist was given a biology question a lab had already solved but not published, and the hypothesis it ranked first matched the confirmed mechanism Can AI systems generate hypotheses that match unpublished experimental discoveries?.
The picture flips when you look at the whole set of ideas instead of any single one. A study of 70+ models across 26,000 open-ended prompts found what it calls an 'Artificial Hivemind': models from different companies independently give strikingly similar answers. They share much of the same training data and go through similar alignment training, so asking several models doesn't buy you as much variety as you'd expect Do different AI models actually produce diverse outputs?. A related critique says AI multiplies *claims* without multiplying *points of view*: a thousand AI-written articles can amount to about one perspective Does AI generate diverse claims or diverse perspectives?. A cultural-theory reading adds that this sameness is harder to see than old mass-media sameness, because every output feels tailored to you Does AI homogenize culture the way mass media did?.
Training is part of the reason. Reinforcement learning rewards whatever strategy scores best, so it pushes models toward a narrow set of winning behaviors. In search agents, RL measurably shrank the range of strategies explored, while fine-tuning on varied human examples kept that range wide Does reinforcement learning squeeze exploration diversity in search agents?. The collapse isn't fixed. It depends on how a model was trained, and different training choices can preserve diversity.
The most useful practical finding is about *where* AI enters the process. In a preregistered writing experiment, letting AI generate ideas made the group's ideas more alike for every writer, while using AI only to refine ideas people already had kept the group's diversity intact. Non-native English speakers brought the most variety, and AI ideation was the one condition that erased that advantage Does AI assistance homogenize or preserve creative diversity?. A second paper is a caution: its ideation study claims AI narrows diversity but never measured diversity, only how many ideas people produced and how detailed they were Does AI assistance actually narrow the diversity of ideas?. So check whether a study measured diversity before you trust its conclusion about it.
What about using several AI agents with different personas to get variety? That helps only up to a point. Multi-agent teams with diverse viewpoints beat a single agent, but only when each member has real domain expertise. Diverse teams without expertise did worse than one competent agent, because the variety turned into noise instead of insight Does cognitive diversity alone improve multi-agent ideation quality?. The takeaway: an AI can surprise you with a single idea, but a whole set of AI ideas tends to come out alike unless humans supply the starting ideas or the agents bring real expertise.
Sources 9 notes
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
When given a question their labs had solved experimentally but not published, the AI platform ranked a hypothesis matching the confirmed mechanism of cf-PICIs hijacking phage tails as its top candidate, suggesting AI can reach established answers independently.
INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.
Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.
AI mass-generates similar flows disguised as personalized outputs, suppressing novelty more deeply than pre-stamped commodities because contextual customization makes homogeneity invisible to individual users. Evidence: independent LLMs converge on similar outputs despite nominal competition.
Show all 9 sources
RL training compresses behavioral diversity in search agents through the same entropy collapse mechanism documented in reasoning—policies converge on narrow reward-maximizing strategies. SFT on diverse demonstrations preserves exploration breadth, suggesting diversity-preservation techniques are essential for RL search scaling.
In a preregistered experiment, AI-generated ideas reduced collective diversity for all writers, while AI that refined existing ideas kept diversity intact. Non-native English speakers contributed more diversity than native speakers, but only AI ideation erased this advantage.
The paper's ideation experiment shows AI help increases idea count and detail, particularly for less experienced writers, but provides no diversity measure to support its conclusion about narrowed diversity.
Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- On Epistemic Diversity in Large Language Models