INQUIRING LINE

Ask five different AI models for startup ideas, and how many genuinely different answers do you actually get?

How much do different LLMs independently converge on similar outputs?

This explores whether separately built AI models, from different companies with different architectures, end up giving the same answers, and what that sameness means for anyone relying on 'more models' to get more variety.


This explores whether different LLMs, built by different teams, independently end up saying the same things, and why that matters. The short answer from the corpus is that they converge a lot. The largest study here looked at more than 70 models across 26,000 open-ended prompts, the kind of question with no single right answer, like 'write a poem about time' or 'suggest a startup idea'. It found what the authors call an 'Artificial Hivemind': models from different labs often produced strikingly similar, sometimes nearly identical, responses Do different AI models actually produce diverse outputs?. The likely causes are overlapping training data and similar alignment procedures. The practical consequence is the surprising part. Asking five different models for ideas may give you less variety than you expect, because ensembles of 'diverse' models are less diverse than their brand names suggest.

The convergence also happens inside a single model. Frontier LLMs show a narrowness: when a question has many valid answers, they keep returning the same small, standard set of answers and reasoning strategies Do frontier LLMs actually explore the full space of valid answers?. Put the two findings together and the problem compounds. Each model narrows to a favorite few answers, and the models share many of the same favorites. A related trap is the belief that fixing a model's settings (temperature zero, a fixed seed) makes its output trustworthy. That only freezes one draw from the model's probabilities and repeats it, so consistency is not reliability Does setting temperature to zero actually make LLM outputs reliable?. Sameness across runs or across models can look like agreement when it is really a shared bias.

Alignment training is a large part of the story. When nine LLMs took standard psychological tests, each kept a stable 'personality' of its own across repeated tests. Underneath, though, all of them converged on the same prosocial, stability-focused profile Do LLMs show reproducible psychological profiles when given standardized tests?. So models differ at the surface but share a common core. That core is strong enough that an aligned model can generate its own training conversations when given only the empty chat template, with no example questions at all Can aligned LLMs generate their own training data?. This is useful for building datasets, but it also shows how much of the 'helpful assistant' pattern has become a template the model reproduces on its own.

Models also tend to fail in the same ways. On constrained optimization tasks, LLMs plateau at about 55–60% constraint satisfaction no matter their architecture, size, or training, and reasoning models do no consistently better Do larger language models solve constrained optimization better?. They tend to recall answers that look like a familiar solution instead of actually running the step-by-step calculation Do large language models actually perform iterative optimization?. One explanation: all these models are, at bottom, machines that predict the next word from probabilities. That fact alone predicts which tasks will be hard for all of them, such as reciting the alphabet backwards, because the correct output is statistically unlikely Can we predict where language models will fail?. Models built on the same basic mechanism and trained on similar data share blind spots as well as habits.

The takeaway you might not have expected: a second model's 'second opinion' is a weak check if both models learned from the same internet and the same alignment recipe. When two models agree, that may only reflect a shared upbringing. Real diversity seems to need outside checks, such as real-world data, external tools, or human judgment, rather than more models from the same family. One example couples LLM-generated candidates with statistical models fitted to actual experimental results Can language models reliably judge their own candidate quality?. The corpus doesn't yet measure how much convergence there is by task type, for instance whether code converges more than fiction. That remains an open question here.


Sources 9 notes

Do different AI models actually produce diverse outputs?

INFINITY-CHAT analyzed 70+ models across 26K open-ended queries and found an "Artificial Hivemind" effect: models independently generate strikingly similar or identical responses due to overlapping training data and alignment procedures, undermining the diversity benefits of model ensembles.

Do frontier LLMs actually explore the full space of valid answers?

Testing across multiple models and domains, researchers found that frontier LLMs often exhibit epistemic narrowness—returning the same valid answers and reasoning strategies repeatedly, even when many alternatives exist. This reduces coverage of the valid answer space users could access.

Does setting temperature to zero actually make LLM outputs reliable?

Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.

Do LLMs show reproducible psychological profiles when given standardized tests?

Nine LLMs given seven psychological instruments in Chinese and English showed stable, model-specific response configurations reproducible across repeated administrations, while all models converged on a shared prosocial and stability-oriented pattern. Structured non-response patterns also distinguished models, suggesting the boundary of what each model treats as applicable is part of its behavioral signature.

Can aligned LLMs generate their own training data?

MAGPIE shows that aligned models like Llama-3-Instruct auto-regressively generate diverse, high-quality instructions when given only pre-query formatting tokens, without prompt engineering. 4M generated pairs matched human-curated datasets in quality and outperformed external sources in downstream fine-tuning.

Show all 9 sources
Do larger language models solve constrained optimization better?

Across constrained-optimization tasks, LLMs converge to ~55–60% constraint satisfaction independent of architecture, parameter count, or training regime. Reasoning models do not systematically outperform standard models, suggesting a fundamental ceiling rather than a scaling gap.

Do large language models actually perform iterative optimization?

Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.

Can we predict where language models will fail?

By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.

Can language models reliably judge their own candidate quality?

LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.