When AI helps people brainstorm, how do we tell if their ideas got more varied, not just more numerous?
What would a valid diversity measure for AI-assisted ideation tasks look like?
This explores what it would take to actually measure whether AI help makes people's ideas more varied or more alike, as opposed to just counting how many ideas they produce.
This explores what a trustworthy test of 'did AI make ideas more or less diverse?' would need to include. The short version from the corpus is that a valid measure has to look at the whole group of ideas rather than any one person's output, and it has to measure differences in perspective rather than differences in wording. A cautionary case comes first. One study concluded that AI help narrows idea diversity, but its experiment only showed that AI increased how many ideas people produced and how detailed they were, especially for less experienced writers. It never measured diversity at all Does AI assistance actually narrow the diversity of ideas?. More ideas and richer ideas are not the same thing as more varied ideas, and a valid measure has to keep those apart.
The first requirement is to measure across people, not within each person. In a preregistered experiment, AI-generated ideas reduced the diversity of the whole pool for every writer, while AI that only refined ideas people already had left that diversity intact Does AI assistance homogenize or preserve creative diversity?. Each person could feel more creative while everyone's output drifted toward the same place, so per-person scores would miss the loss entirely. The same study turned up a detail a per-person metric would also miss: non-native English speakers added more diversity than native speakers, and only AI ideation erased that advantage. So a good measure also asks whose distinctive contributions disappear. The same logic applies one level up, to the models themselves. More than 70 different models tend to give strikingly similar answers to open-ended prompts Do different AI models actually produce diverse outputs?. Switching between AI tools doesn't create a diverse baseline, and the right comparison is ideas produced with no AI at all.
The second requirement is to measure at the right level of difference. Variation in wording can rise while variation in viewpoint stays flat: AI can produce many well-formed claims that all come from roughly one point of view Does AI generate diverse claims or diverse perspectives?. Which level you measure can even flip the result. Preference tuning (RLHF) reduces word-and-syntax variety in code but increases it in creative writing Does preference tuning always reduce diversity the same way?. Research on model reasoning points to a more meaningful unit, which is the underlying approach or strategy. Spreading effort across different high-level approaches beats generating many variations of one approach Can abstractions guide exploration better than depth alone?, and setting up reasoning as a dialogue between distinct viewpoints produces more genuinely different problem-solving paths Can dialogue format help models reason more diversely?. For ideation, that suggests grouping ideas by the approach behind them before counting how varied they are.
The third requirement is to pair diversity with quality and follow it over time. A diversity score with nothing alongside it is easy to game: teams with varied thinking styles but no real expertise did worse than a single competent agent, because the variety produced confusion instead of insight Does cognitive diversity alone improve multi-agent ideation quality?. Similarly, LLM research ideas were rated more novel than experts' ideas but slightly less feasible Do language models generate more novel research ideas than experts?, so novelty and usefulness need separate scores. Time matters too. In model training, variety in solutions narrows gradually over repeated rounds unless something actively counteracts it Do critique models improve diversity during training itself?. If people use AI ideation over weeks, a single snapshot could miss a slow drift toward sameness.
Put together, a valid measure would do five things. It would compare the group's ideas against a no-AI baseline. It would group ideas by the perspective behind them, not by their wording. It would separate where the AI comes in, generating ideas versus refining them. It would track which people's distinctive contributions shrink. And it would report quality next to diversity, repeated over time. What may be surprising is that the most influential claim in this area, that AI narrows ideas, comes in at least one case from a study that never measured diversity. The finding that holds up best is more specific: the harm comes from letting AI generate the ideas, while using it to refine ideas leaves diversity intact.
Sources 10 notes
The paper's ideation experiment shows AI help increases idea count and detail, particularly for less experienced writers, but provides no diversity measure to support its conclusion about narrowed diversity.
In a preregistered experiment, AI-generated ideas reduced collective diversity for all writers, while AI that refined existing ideas kept diversity intact. Non-native English speakers contributed more diversity than native speakers, but only AI ideation erased this advantage.
INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.
Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.
RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.
Show all 10 sources
RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.
DialogueReason, which structures a single model's internal reasoning as dialogue between distinct agents in separate scenes, overcomes monologue reasoning's fixed-strategy and fragmented-attention weaknesses, especially on tasks requiring multiple problem-solving approaches.
Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.
A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.
Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- NoveltyBench: Evaluating Language Models for Humanlike Diversity
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers