INQUIRING LINE

When AI helps people brainstorm, how do we tell if their ideas got more varied, not just more numerous?

What would a valid diversity measure for AI-assisted ideation tasks look like?

This explores what it would take to actually measure whether AI help makes people's ideas more varied or more alike, as opposed to just counting how many ideas they produce.


This explores what a trustworthy test of 'did AI make ideas more or less diverse?' would need to include. The short version from the corpus is that a valid measure has to look at the whole group of ideas rather than any one person's output, and it has to measure differences in perspective rather than differences in wording. A cautionary case comes first. One study concluded that AI help narrows idea diversity, but its experiment only showed that AI increased how many ideas people produced and how detailed they were, especially for less experienced writers. It never measured diversity at all Does AI assistance actually narrow the diversity of ideas?. More ideas and richer ideas are not the same thing as more varied ideas, and a valid measure has to keep those apart.

The first requirement is to measure across people, not within each person. In a preregistered experiment, AI-generated ideas reduced the diversity of the whole pool for every writer, while AI that only refined ideas people already had left that diversity intact Does AI assistance homogenize or preserve creative diversity?. Each person could feel more creative while everyone's output drifted toward the same place, so per-person scores would miss the loss entirely. The same study turned up a detail a per-person metric would also miss: non-native English speakers added more diversity than native speakers, and only AI ideation erased that advantage. So a good measure also asks whose distinctive contributions disappear. The same logic applies one level up, to the models themselves. More than 70 different models tend to give strikingly similar answers to open-ended prompts Do different AI models actually produce diverse outputs?. Switching between AI tools doesn't create a diverse baseline, and the right comparison is ideas produced with no AI at all.

The second requirement is to measure at the right level of difference. Variation in wording can rise while variation in viewpoint stays flat: AI can produce many well-formed claims that all come from roughly one point of view Does AI generate diverse claims or diverse perspectives?. Which level you measure can even flip the result. Preference tuning (RLHF) reduces word-and-syntax variety in code but increases it in creative writing Does preference tuning always reduce diversity the same way?. Research on model reasoning points to a more meaningful unit, which is the underlying approach or strategy. Spreading effort across different high-level approaches beats generating many variations of one approach Can abstractions guide exploration better than depth alone?, and setting up reasoning as a dialogue between distinct viewpoints produces more genuinely different problem-solving paths Can dialogue format help models reason more diversely?. For ideation, that suggests grouping ideas by the approach behind them before counting how varied they are.

The third requirement is to pair diversity with quality and follow it over time. A diversity score with nothing alongside it is easy to game: teams with varied thinking styles but no real expertise did worse than a single competent agent, because the variety produced confusion instead of insight Does cognitive diversity alone improve multi-agent ideation quality?. Similarly, LLM research ideas were rated more novel than experts' ideas but slightly less feasible Do language models generate more novel research ideas than experts?, so novelty and usefulness need separate scores. Time matters too. In model training, variety in solutions narrows gradually over repeated rounds unless something actively counteracts it Do critique models improve diversity during training itself?. If people use AI ideation over weeks, a single snapshot could miss a slow drift toward sameness.

Put together, a valid measure would do five things. It would compare the group's ideas against a no-AI baseline. It would group ideas by the perspective behind them, not by their wording. It would separate where the AI comes in, generating ideas versus refining them. It would track which people's distinctive contributions shrink. And it would report quality next to diversity, repeated over time. What may be surprising is that the most influential claim in this area, that AI narrows ideas, comes in at least one case from a study that never measured diversity. The finding that holds up best is more specific: the harm comes from letting AI generate the ideas, while using it to refine ideas leaves diversity intact.


Sources 10 notes

Does AI assistance actually narrow the diversity of ideas?

The paper's ideation experiment shows AI help increases idea count and detail, particularly for less experienced writers, but provides no diversity measure to support its conclusion about narrowed diversity.

Does AI assistance homogenize or preserve creative diversity?

In a preregistered experiment, AI-generated ideas reduced collective diversity for all writers, while AI that refined existing ideas kept diversity intact. Non-native English speakers contributed more diversity than native speakers, but only AI ideation erased this advantage.

Do different AI models actually produce diverse outputs?

INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.

Does AI generate diverse claims or diverse perspectives?

Large language models generate numerous well-formed claims by following probabilistic patterns in training data, not by exploring competing argumentative positions. This produces volume without perspectival diversity—a thousand AI articles often represent approximately one viewpoint.

Does preference tuning always reduce diversity the same way?

RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.

Show all 10 sources
Can abstractions guide exploration better than depth alone?

RLAD jointly trains abstraction and solution generators, showing that allocating test-time compute to diverse abstractions outperforms parallel solution sampling at large budgets. Abstractions create structured breadth-first exploration that prevents the underthinking failure mode of depth-only reasoning chains.

Can dialogue format help models reason more diversely?

DialogueReason, which structures a single model's internal reasoning as dialogue between distinct agents in separate scenes, overcomes monologue reasoning's fixed-strategy and fragmented-attention weaknesses, especially on tasks requiring multiple problem-solving approaches.

Does cognitive diversity alone improve multi-agent ideation quality?

Multi-agent teams substantially outperform solo ideation, but only when members possess genuine senior knowledge. Diverse teams without expertise underperform even a single competent agent, because cognitive stimulation without expertise triggers process losses instead of insight.

Do language models generate more novel research ideas than experts?

A statistically significant study of 100+ NLP researchers found LLM-generated ideas rated as more novel than human expert ideas (p<0.05), though slightly lower on feasibility. Expert knowledge constrains novelty, while LLMs explore wider conceptual combinations.

Do critique models improve diversity during training itself?

Step-level critique in the training loop counteracts tail narrowing and maintains solution diversity across self-training iterations. This training-time benefit—preventing premature convergence—is more fundamental than test-time accuracy gains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.