Six misconceptions about large language models: A minimal model and diagnostic taxonomy
Abstract Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories—intuitive, informal explanatory models that guide attitudes and actions. Deflationary slogans (“just autocomplete,” “stochastic parrots,” and “average of the internet”) and anthropomorphic framings (“emergent agents” and “protominds”) each capture genuine features of current systems but mistake those features for the whole. This Perspective proposes a minimal working model of LLM-based systems centered on four distinctions: between pretraining and deployed systems; between the learned distribution and particular samples; among parametric, contextual, and external memory; and between task competence and agency. The model is used to diagnose six misconceptions about LLMs: nexttoken prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. For each, the analysis identifies what the misconception gets right, which distinctions it conflates, and what follows for capability evaluation, system design, and governance.
Introduction. Few debates about scientific practice, education, or creative work now go without invoking large language models (LLMs). Journalists, critics, and researchers use ready-made slogans to declare what these systems “really” are: “glorified autocomplete” or “just next-token predictors”; “stochastic parrots”; “lossy text-compression algorithms” or “a blurry JPEG of all the text on the Web”; “bullshit generators”; “high-tech parlor tricks”; or “the average of the internet, edited for tone” (1–3). Others cast them in more anthropomorphic terms: as “superhuman reasoners,” “proto-agents,” or instances of the “wisdom of the silicon crowd” (4– 7). These slogans can morph into folk theories: intuitive, informal explanatory models— often partial and tacit—used to make sense of technological systems and guide action toward them (8, 9). “Folk” refers to the mode of explanation rather than the speaker’s sophistication; indeed, informal and technical accounts can coexist within researchers, policymakers, and lay users alike.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations?- Why do both deflationary and anthropomorphic framings of LLMs persist in research?
- Can LLMs infer situational context the way humans do pragmatically?
- How do fixed pragmatic templates prevent models from understanding context?
- What happens when LLMs analyze literary irony that relies on understatement?
- What should we call errors in LLM outputs when hallucination does not apply?
- How does LLM hallucination risk manifest in knowledge graph construction?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Where do LLMs succeed at generation but struggle with evaluation?
- Why do LLM personas struggle with specificity in specialized domains like law?
- Why do LLMs generate ideas that sound novel but fail during execution?
- What specific execution barriers do LLM ideas encounter most frequently?
- Can output-layer corrections fix fundamental cultural representation deficits in LLMs?
- Why do LLM explanations feel authoritative even when alignment with the model fails?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- What makes a problem instance unfamiliar to a language model?
- Why do NLP benchmarks exclude ambiguous instances from evaluation?