Do AI chatbots narrow the range of answers they give, and can we measure that shrinkage instead of just arguing it?
Can measuring answer-space collapse show LLM narrowing in practice?
This explores whether we can catch LLMs narrowing what they say in practice, by measuring how much the range of possible answers shrinks, rather than just arguing that it happens in theory.
This explores whether the idea that LLMs narrow human expression can be tested by measuring how much their range of answers shrinks. The short version: the collection makes a strong case that narrowing happens, but it has little on directly measuring answer-space collapse. Most of what it offers points at where to look and why ordinary tests miss it.
The main case for narrowing is in Do large language models narrow human expression and thought?. Models reflect a skewed slice of human experience from their training data. Because millions of people use the same few models, that skew adds up into convergence. Co-writing studies show the effect in action: people take on the model's stances and framings without noticing. That's evidence of narrowing in *people*, though, not a measure of how much the model's own answers have collapsed.
For a way to predict where collapse should show up, look at Can we predict where language models will fail?. If you treat an LLM as a machine that predicts the most probable next word, you can predict that tasks with unlikely correct answers will be hard even when they're logically simple, like reciting the alphabet backwards. Narrowing is the same pull toward the probable. So this work gives you a model for predicting *which* answers will get squeezed out before you measure anything. The same pull shows up in other forms. Models fall back on familiar meanings when the logic is unfamiliar (Do large language models reason symbolically or semantically?). They emit memorized, template-like values instead of actually working through iterative calculations (Do large language models actually perform iterative optimization?). And very different models land on the same ~55–60% ceiling on constraint problems (Do larger language models solve constrained optimization better?). That last one is a form of narrowing across models: different architectures, sizes and training methods all end up in the same place.
The most surprising point for measurement is that standard benchmarks may be built to hide narrowing. Do standard NLP benchmarks hide LLM ambiguity failures? shows that benchmarks drop exactly the examples where human annotators disagree, which are the cases with more than one legitimate answer. When those cases are put back, accuracy falls from about 90% to 32%. If you want to see a model collapse many valid answers into one, you need the messy, disputed questions that current evaluations throw out.
One lead points toward watching collapse as it happens. In diffusion LLMs, which refine a whole response at once instead of writing it left to right, the model's confidence in its answer settles early while its reasoning keeps changing (Can reasoning and answers be generated separately in language models?). That paper uses the effect to save compute, not to study narrowing. Still, it suggests that the moment an answer locks in can be observed and timed, which is the kind of tool a real measure of answer-space collapse would need. The collection hasn't yet joined these pieces into a direct measurement.
Sources 7 notes
LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.
By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.
When semantic content is decoupled from reasoning tasks, LLM performance collapses even with correct rules in context. Models rely on parametric commonsense and token associations rather than formal logical manipulation, constraining reasoning to training distribution semantics.
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
Across constrained-optimization tasks, LLMs converge to ~55–60% constraint satisfaction independent of architecture, parameter count, or training regime. Reasoning models do not systematically outperform standard models, suggesting a fundamental ceiling rather than a scaling gap.
Show all 7 sources
By filtering out examples where annotators disagree, benchmarks remove test cases that would reveal LLM failures at ambiguity recognition. Research using ambiguous examples shows a 32% vs. 90% accuracy gap invisible to standard evaluation.
ICE shows that bidirectional attention in diffusion LLMs enables in-place prompting—embedding reasoning directly in masked positions refined alongside answers. Answer confidence converges early while reasoning continues refining, allowing early-exit mechanisms to cut compute by 50% while maintaining accuracy.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Can Large Language Models Reason and Optimize Under Constraints?
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
- Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
- Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models