Do LLM improvements reflect reasoning gains or corpus shifts?
When large language models improve on previously failed tasks, does this show they've learned to reason better, or are they simply reflecting changes in human-written text they train on? Understanding this matters for assessing what LLMs actually know.
The authors argue that an LLM's improvement on a question it once got wrong reflects "the change in the body of human written discourse on which they can draw," not a gain in reasoning. Their case is Colin Fraser's "Dumb Monty Hall problem": ChatGPT failed it in 2023, and LLMs were succeeding "a year or so later." A quoted ChatGPT-4o answer rejects a vaccine–autism link, and the authors read that consensus as the corpus speaking. Their verdict is that "the 'intelligence', we argue, is in the humans not the LLMs."
The mechanism starts with how science makes knowledge. The authors draw on Collins's case studies and a 2014 study of gravitational-wave physicists who set aside a "fringe science" paper, a decision made "largely as a result of tacit knowledge built up in social discourse, mostly spoken discourse, within closed groups of experts." New ideas enter with almost no written trace, which the authors call "a good thing because most new ideas are useless." LLMs rest "at base, on statistical relationships in the corpus of written discourse," so they can separate fringe from path-breaking ideas only "by mimicking human judgements once they have been represented in the published, or at least formally written, corpus." Settled positions can be held in place by alignment, which the authors liken to "The Ministry of Truth" in the LLM world.
Against the nearest notes, this excerpt gives a mechanism for the claim in Do classical knowledge definitions apply to AI systems?. That note argues that knowledge without a human knower breaks the classical definition. This excerpt says where the human side does its work: in the spoken, bounded phase of science before anything is written. It also complements Do large language models narrow human expression and thought?. That paper locates narrowing in the middle of the distribution, where dominant styles are over-represented. This excerpt locates a gap at the edge, where new ideas should enter. The same blind spot appears in Can AI distinguish which differences actually matter?, here placed in the timing of the corpus rather than in the act of observation.
The excerpt does not establish much of what its argument needs. The new Monty Hall prompt, the human panel's results and the 18-model table are described but not included, and the two figures are referred to without being reproduced. The only LLM response shown is a single April 2024 transcript. The claims that LLMs will "soon" re-align with humans, and that the Dumb Monty Hall change came from corpus shifts, are argued rather than tested against any training data. The excerpt also says it does not deny "model-side improvements arising from architecture, scale, or training procedures," so the corpus account is one explanation among several. The overshadowing section named in the abstract is absent. At the strength the evidence allows, the corpus explanation is a testable hypothesis: track answers over time against the corpus changes behind them. The authors propose that test, but the excerpt does not run it.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can AI systems reliably guide voters without introducing political bias?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do classical knowledge definitions apply to AI systems?
Classical definitions of knowledge assume truth-correspondence and a human knower. Do these assumptions hold for LLMs and distributed neural knowledge systems, or do they need fundamental revision?
the classical-definition argument; this excerpt names where the human side of knowledge formation happens, in pre-written spoken discourse
-
Do large language models narrow human expression and thought?
Explores whether LLMs homogenize how people write, think, and reason by reflecting narrow training distributions and subtly shifting user preferences toward model outputs.
locates narrowing in the middle of the distribution; this excerpt locates the missing new idea at its edge
-
Can AI distinguish which differences actually matter?
Explores whether AI systems can perform the qualitative judgment that experts use to select relevant observations. Matters because confusing AI outputs with expert observation leads users to trust pattern-matching as if it were reasoning about what's important.
same blind spot, qualitative selection of relevance, here placed in corpus timing rather than observation
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
- Word Meanings in Transformer Language Models
- LLMs can implicitly learn from mistakes in-context
- Available but Unclaimed: An Empirical Study of Human-AI Synergy
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- The Thin Line Between Comprehension and Persuasion in LLMs
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
Original note title
LLM answers track shifts in human written discourse rather than reasoning gains — the authors argue the intelligence sits in the humans