Can language models learn meaning without engaging the world?
Explores whether LLMs prove that meaning emerges from relational structure alone, independent of embodied experience or external reference. Tests structuralist theory empirically.
"Computational Structuralism: Toward a Formal Theory of Meaning in the Age of Digital Intelligence" (2026) proposes a synthesis of deep learning, information theory, and French structuralism to interpret LLM success. The core argument: LLMs demonstrate that transformations over relational structure are sufficient for generating culturally and situationally specific discourse, and that such structure can be inductively derived from discourse traces alone — phenomenal or embodied engagement with the world is not a necessary condition.
The framework retraces the lineage from Saussure (language as a system of differences, meanings defined relationally) through Levi-Strauss (extending structural analysis to culture broadly, binary oppositions as compression of complexity) to Bourdieu (habitus as transposable classification schemas operating in continuous social space). LLMs trained on web text learn not just grammar but the structure of culturally situated linguistic action — which voices make which statements in response to which situations, and how audiences respond.
Key theoretical moves:
- LLMs operationalize Saussure's concept of langue — not the set of all valid statements, but the system that can interpret and generate all valid statements
- Language modeling is equivalent to text compression: removing redundancies by replacing them with generative principles. The same statistical dependencies that inform prediction compose the compressed model
- The framework privileges sufficiency over necessity — LLMs drawing on the same operations as humans is not claimed, but one way to achieve fluent natural language is now formally demonstrated
- Mechanistic interpretability offers the possibility of reverse-engineering these latent structures, answering structuralist questions (how are ideologies composed from simpler features?) with empirical methods
This challenges both sides of the grounding debate: it validates the structuralist intuition that relational form can carry meaning without referential content, while simultaneously showing that what LLMs learn is not "pure language" but socially and culturally situated discourse patterns. The concern from Can language models learn meaning from text patterns alone? (Bender & Koller) is not refuted but reframed — what counts as "sufficient" for meaning generation may not require what's necessary for meaning understanding.
Connects to Does semantic grounding in language models come in degrees? — computational structuralism explains why functional grounding succeeds: the relational structure of discourse is compressible and learnable. The question is whether this constitutes meaning or merely its simulation.
Inquiring lines that read this note 130
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do language models hallucinate and how can we prevent it?- Can fixing hallucination address AI's structural epistemic problem?
- How does interleaving reasoning with action prevent hallucination in language models?
- Can secondary orality exist without any embodied human participant at all?
- Can we develop competent reading practices for disembodied orality?
- How does training data preserve communicative event structure without the actual events?
- Can statistical learning from language alone capture all aspects of cultural competence?
- What makes a relational act different from just moving content around?
- How does monological training on text differ from dialogical training in conversation?
- Can a virtual instance be individuated from its conversational context?
- Can knowledge flow without an embodied carrier transmitting it?
- How does enactive theory define language differently than computational linguistics?
- Can linguistic agency exist without embodiment and real-world participation?
- How do low-dimensional representation structures entangle multiple cultures together?
- What makes linguistic agency impossible for systems without embodiment?
- Does selective suppression of linguistic relations enable human meaning-making?
- Does embodiment matter for genuine linguistic agency?
- What role does joint attention play in how humans learn language meaning?
- Can language meaning emerge without joint attention and shared embodied interaction?
- How does embodiment affect whether LLMs can participate in Wittgensteinian language games?
- Can understanding language happen entirely within a language system alone?
- Can functional behavior alone capture what makes something a genuine belief?
- What does embodiment and precariousness mean for linguistic agency?
- Why does conceptual priming alone fail to produce consciousness claims?
- What makes the Extended Mind thesis incompatible with internalism?
- What role does the biological substrate play in human relational identity?
- Why does joint attention matter for acquiring linguistic meaning?
- Can statistical learning from text replace embodied cultural experience?
- Does language convey meaning purely through relational structure without external grounding?
- Does computational functionalism require an experiencing agent to ground symbols?
- Do language models raise validity claims in the Habermasian sense?
- Can a relational entity bear psychological properties the way Chalmers claims?
- Does functional grounding through discourse patterns count as genuine semantic meaning?
- Can LLMs participate meaningfully in discourse without consciousness or understanding?
- What structural limits prevent LLMs from abstracting moral principles?
- Can LLMs predict social norms without deep integration into linguistic practices?
- How does Wittgenstein's language games explain social grounding in LLMs?
- Can you separate grammatical competence from rhetorical commitment in language systems?
- What makes relational structure sufficient for generating contextually appropriate discourse?
- Does embodiment and interaction matter for linguistic competence beyond pattern learning?
- What role does language play as a cognitive scaffold versus communication tool?
- How does subject-predicate distinction emerge from formal linguistic analysis?
- Can pragmatic competence emerge from text exposure alone without interactive grounding?
- Can pragmatic competence emerge from text exposure without interactive grounding?
- Can LLMs infer situational context the way humans do pragmatically?
- How do humans learn language through communication differently than LLM text prediction?
- How does syntactic encoding relate to semantic feature representation?
- How does semantic grounding differ between human minds and language models?
- Can language models reason without relying on learned semantic patterns?
- Can LLMs improve at metaphor if they handle decoupled semantics better?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why do language models fail at implicit discourse relations while handling explicit connectives?
- Why do explicit discourse connectives help LLMs but implicit relations cause failures?
- How does the symbol grounding problem apply to artificial language systems?
- Can LLMs infer implicit meaning without surface linguistic markers?
- Can language models develop world models that ground meaning in causal reality?
- How do internal representations compare to human cognitive structures?
- Do language models actually learn linguistic structure or just surface statistics?
- Do metaphors work by decoupling meaning from linguistic associations?
- Can LLMs identify implicit metaphoric mappings that require pragmatic inference?
- Can LLM semantic representations exist without causally influencing their generation output?
- Do language models encode deep syntactic structure or only surface-level patterns?
- Why does LLM compression eliminate causal grounding in conceptual representations?
- Why do explicit discourse connectives work when implicit relations fail?
- Does DPO training with coreference chains teach spontaneous convention formation?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- How does bidirectional entailment distinguish semantic equivalence from token similarity?
- Why do language models reproduce human EPA structure despite different architecture?
- Can language models learn internal world models without explicit environment specifications?
- Can LLMs reason through semantics without understanding causal mechanisms?
- Do language models need words to think or just latent structure?
- Do LLMs learn abstract grammar or culturally situated discourse patterns instead?
- Can language acquisition analogies mislead us about how models actually learn?
- Why does frame-activation matter more than word-by-word composition?
- Can frame semantics explain why context matters more than word similarity?
- How do static embeddings and contextualized representations divide semantic labor?
- Can readers detect meaning through resonance patterns alone without knowing authorial intent?
- Where does the meaning actually originate in reader-detected resonance across language?
- Why does training data saliency distort how models judge meaning?
- Can implicit linguistic information ever be reliably learned from training data?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- How does mechanistic interpretability reveal ideological structures in language model weights?
- How do world models decompose between representation of facts versus generative mechanisms?
- What makes internal embeddings useful as multimodal input for language model training?
- Why do users attribute consciousness to language models in practice?
- Do language models learn surface patterns instead of underlying linguistic principles?
- Can large language models understand language without embodied grounding systems?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- What architectural changes would let language models develop genuine functional competence?
- What distinguishes surface cues from structural meaning in language understanding?
- What structural properties of language models make fabrication inevitable?
- Can encoder models match human conceptual structure better than larger language models?
- What distinguishes surface generalizations from true linguistic generalizations?
- Why do surface generalizations fail on unusual syntactic structures?
- Can formal language pretraining address surface generalization without learning true linguistic structure?
- Why do explicit linguistic markers override semantic computation in models?
- How do language models transmit traits through semantically unrelated data?
- What distinguishes real understanding from superficial pattern matching?
- Are static embeddings analogous to the formal linguistic competence layer?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- How does co-occurrence statistics alone produce hierarchical concept organization?
- Why do language models need external temporal signals at all?
- How do world models create indirect causal grounding without physical environment contact?
- Can functional semantic grounding substitute for true causal grounding?
- Can external actions provide causal necessity that language models lack?
- Can LLMs develop genuine understanding without embodied experience?
- Does LLM vocabulary become the cultural lexicon for how we think?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- How do speech encoders learn articulatory physics without phonetic labels?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- Why does gradient descent discover compositional structure without explicit pressure?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- How should we rethink the symbolism versus connectionism debate in light of LLMs?
- How do semantic features in representations become steerable task-specific directions?
- What makes a new representational primitive valuable enough to justify its representational cost?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Mechanistic Indicators of Understanding in Large Language Models
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- Semantic Structure in Large Language Model Embeddings
- What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
Original note title
LLMs operationalize Saussures langue — fully relational models with no external referents suffice to generate contextually appropriate discourse