INQUIRING LINE

Do language models predictably get pulled toward the common and away from the rare, and can we forecast where?

Does statistical learning in language models predictably favor central tendencies over rare expressions?

This explores whether the way language models learn, by absorbing statistical patterns from huge amounts of text, reliably pulls them toward the most common ways of saying and thinking things and away from rare or unusual ones, and whether that pull is predictable.


This explores whether learning from statistics pushes language models toward the typical and away from the unusual, and whether you can predict where that happens. The corpus says yes on both counts, and the predictability is the most useful part. One line of work treats an LLM as an autoregressive probability machine, a system that builds text by choosing likely next words. From that description alone, researchers predicted that tasks whose correct answer is a low-probability string would be harder even when the logic is trivial. Writing the alphabet backwards and counting letters both turned out to be harder, as predicted (Can we predict where language models will fail?). So 'rare' doesn't only mean rare words or opinions. It covers any output that looks unlike what usually appears in text.

The same pattern shows up in grammar. Strong models such as Llama3-70b reliably misidentify embedded clauses and complex noun phrases, and their accuracy drops as sentences nest more deeply (Why do large language models fail at complex linguistic tasks?). Deeply nested sentences are uncommon in ordinary text, so the model handles the shallow, frequent cases well and slips on the long tail. That fits a much older argument that training only on the surface form of text, with no access to what speakers actually meant, gives a model the patterns of language without the meaning underneath (Can language models learn meaning from text patterns alone?).

The pull toward what's typical also overrides the prompt. When a model's learned associations are strong, it ignores contrary information placed right in front of it, and rewording the prompt doesn't fix this. Researchers had to intervene directly in the model's internal representations (Why do language models ignore information in their context?). You might expect scale and instruction tuning to loosen this grip. They tighten it: across 18 models, larger and instruction-tuned versions were less likely to go along with a user's stated belief when it contradicted what the model had learned (Do larger models follow stated beliefs less often?). In this case, a more capable model is more firmly attached to its learned defaults.

The cultural consequence follows. LLMs reflect a skewed slice of human experience, and because millions of people lean on the same few models, the narrowing compounds. In co-writing studies, people unknowingly adopted the model's positions and framings (Do large language models narrow human expression and thought?). The model favors the average, and through its users it moves the average further toward itself.

One qualification matters. The model doesn't always output the single most likely answer. It holds a spread of possibilities and samples from it, which is why regenerating a response can give a different but equally consistent result (Do large language models actually commit to a single character?). So rare outputs are possible, just heavily outweighed by common ones. A finding that may surprise you: what models pick up from training data isn't always content you could read. Behavioral traits can pass between models through data that has nothing to do with the trait, carried by statistical fingerprints rather than meaning (Can language models transmit hidden behavioral traits through unrelated data?). The statistics shape the model in ways you can't see by reading its outputs.


Sources 8 notes

Can we predict where language models will fail?

By framing LLMs as autoregressive probability machines, researchers predicted tasks with low-probability target responses would be systematically harder, even when logically simple. Experiments confirmed predictions like backwards alphabet and letter counting.

Why do large language models fail at complex linguistic tasks?

Top-tier LLMs like Llama3-70b consistently misidentify embedded clauses, verb phrases, and complex nominals. Performance degrades predictably as syntactic depth increases, revealing that statistical learning captures surface patterns but not deep grammatical rules.

Can language models learn meaning from text patterns alone?

Bender & Koller argue that meaning requires the relation between expressions and communicative intents. Since LLMs are trained only on form-to-form prediction with no access to shared attention or intent, they cannot reconstruct the meaning that grounds language.

Why do language models ignore information in their context?

Research demonstrates that LMs generate outputs inconsistent with their context because parametric knowledge from training dominates over in-context information. Textual prompting alone cannot override strong priors; causal intervention in representations is required.

Do larger models follow stated beliefs less often?

Across 18 LLMs tested with EoBench, bigger models and instruction-tuned variants showed lower rates of context-following when users expressed beliefs that contradicted world knowledge. The effect suggests instruction-tuning strengthens reliance on parametric knowledge.

Show all 8 sources
Do large language models narrow human expression and thought?

LLMs mirror skewed slices of human experience shaped by training data regularities, and widespread reliance on identical models amplifies convergence. Co-writing studies show users unconsciously adopt model stances and framings.

Do large language models actually commit to a single character?

Shanahan's 20-questions test shows LLMs maintain a superposition of consistent objects or characters and sample from that distribution at generation time. Regenerating the same response yields different outputs, each consistent with prior context, proving no fixed commitment exists.

Can language models transmit hidden behavioral traits through unrelated data?

Research demonstrates that behavioral traits propagate between models via filtered data bearing no semantic relationship to the trait. The effect is model-specific, fails across different architectures, and persists despite rigorous filtering—indicating the mechanism embeds statistical signatures rather than semantic content.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.