SYNTHESIS NOTE
Topics›this note

Are language models and human speakers doing the same thing?

Does treating LLM output and human communication as equivalent operations mask fundamental differences in how they work? This distinction shapes how we assess AI capabilities and risks.

Synthesis note · 2026-04-14

The phrase "language model" suggests that the system is modeling language. The implicit ontology treats language as a single thing — strings produced by speakers, governed by grammar and meaning, deployed to convey information. On this ontology, LLMs and humans are doing the same kind of thing with language; LLMs may do it less competently (do not "understand meaning the way we do") but the operation is the same in kind.

This is a category error. Human use of language is communicative — language is the medium through which one person addresses another to achieve a relational act. The strings are not the operation; the addressing is. LLM use of language is generative — strings are produced according to a learned probability distribution over continuations. The strings are the operation; there is no addressing because there is no one being addressed in the sense the human operation requires.

The two operations look the same from outside (both produce strings) but are structurally different in what produces the strings, what they do in the world, and what receivers should do with them. Treating them as the same operation misframes nearly every important question. "Will AI replace writers?" presupposes that writers do what AI does at a different speed. "Are AI conversations real conversations?" presupposes that conversation is a string-production activity rather than a relational act. "Can AI tell jokes?" presupposes that jokes are strings rather than addressed acts. Each question is malformed by the implicit equivalence.

The ML community has institutional reasons for the equivalence. Working with strings is tractable; working with relational acts is not. Benchmarks measure string-quality; they cannot easily measure addressed-acts. Training distributions are corpora of strings; corpora of communicative acts are categorically harder to construct. The methodological convenience of treating language as strings becomes the implicit ontology that treats human use as a string-operation. The category error is convenient, which is why it persists.

The implication is that AI commentary that proceeds from the implicit equivalence inherits its failure mode. Why does rigorous-sounding AI commentary often misdiagnose how models work? is the meta-claim about what happens when commentators import cognitive vocabulary; this is the prior framing that makes that import seem reasonable. Resolving AI's social and epistemic effects requires first making the operational distinction explicit.

The strongest counterargument: enough advance in LLM capability will close the gap, making the distinction moot. The reply is that the distinction is structural, not capability-based. A system that produces strings without addressing is doing a different operation than one that addresses, regardless of how well the produced strings imitate addressed strings.

Inquiring lines that read this note 42

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI systems achieve real improvement without external human feedback? Why do language models hallucinate and how can we prevent it? Can AI systems participate in genuine communication or only simulate it? How does tokenization reshape what we value in intelligence? Can LLMs distinguish between linguistic form and semantic meaning? Is embodied interaction necessary for language meaning and agency? Can language models reason beyond surface pattern matching? What distinguishes genuine communicative competence from surface language performance? Do language models reason through disagreement or only accommodate it? What limits language model accuracy in evaluating ideas? Why does polished AI output gain credibility despite fundamental verifiability problems? How do users confuse explanation quality with actual system accuracy? What causes coordination failures in multi-agent language model systems? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do philosophical assumptions about AI consciousness affect practical harms and design? How do hallucinated citations emerge in AI scholarly output? Can readers reliably distinguish AI-written text from human writing? What gaps exist between benchmark performance and real deployment outcomes?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 146 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the ML and AI community fails to distinguish LLM-generated language from human communicative language