Can we fairly say an AI 'believes' or 'wants' something, without first solving whether it actually feels anything?
Can non-phenomenal mental states like belief apply to LLMs functionally?
This explores whether we can fairly say an LLM "believes" or "wants" things in a working, functional sense, without first having to settle whether it has any inner experience or consciousness.
This explores whether we can fairly say an LLM "believes" or "wants" things in a working, functional sense, without first having to decide whether it has any inner experience. The corpus leans toward yes, with conditions. Philosophers have taken the consciousness question off the table on purpose. Chalmers' quasi-interpretivism says a system has belief-like states if treating it as having beliefs reliably explains and predicts what it does, and that holds whether or not anything is felt Can we describe LLM beliefs without assuming consciousness?. A related position, "modest inflationism," argues that the usual debunking moves don't hold up: "it only predicts tokens" and "its outputs are too fragile to count" both assume the answer they're trying to prove. The proposal is to treat LLMs the way we treat animals. We readily say a dog believes its owner is home without claiming to know what being a dog feels like Can we defend modest mental attributions to large language models?.
What makes the separation workable is that the hard skeptical arguments mostly target consciousness, not belief. Hoel's formal argument that no testable theory of consciousness can apply to LLMs Can any falsifiable theory of consciousness apply to LLMs? and the view that consciousness requires sharing a physical world with us Can disembodied language models ever qualify as conscious? are strong claims. Neither one, however, rules out functional belief. The surprising consequence is that "no consciousness" and "has beliefs" can both be true together.
The empirical work is where the functional view gets tested, and the results are mixed in a useful way. A functional belief should guide action. Yet in Trust Game experiments, LLMs state plausible beliefs for a persona and then act against them. Adding more explicit context made the gap worse, not better Why do LLMs fail to act on their stated beliefs?. On the functionalist's own standard, that's a problem: a "belief" that doesn't drive behavior is closer to a description than a belief. Chalmers sees the same limit and draws the line at simple internal states, since the interpretation stops working once you get to relational or normative acts like making a promise or an assertion Can we describe LLM beliefs without assuming consciousness?.
There's a second test: can models track beliefs, both other people's and their own? They do well with fixed states, like a persuader's unchanging goal, but fall behind humans when a mind changes partway through a conversation Can language models track how minds change during persuasion?. In open-ended settings they often rely on surface patterns instead of actually modeling another person. Hybrid systems that force explicit belief tracking do better Do large language models genuinely simulate mental states?. Measurable signs of reasoning about others do show up in economic games, but they vary a lot by provider and model size Do LLMs use inferred beliefs to adapt their game strategies?. So "do LLMs have beliefs?" may not have one answer. It may depend on the model.
The takeaway you might not have expected: once consciousness is set aside, the real question about LLM belief is an engineering one, namely whether stated beliefs are wired to action and update when evidence arrives. Researchers simulating society argue that current agents are "stuck in behaviorism" and need explicit belief networks to get there Can language models simulate belief change in people?. There's also a warning about the language itself. Talking about LLMs as believers can quietly change how we describe human minds too How does LLM vocabulary spread beliefs about human thinking?.
Sources 10 notes
Chalmers introduces quasi-interpretivism to ascribe belief-like states to LLMs based on behavioral interpretability without committing to phenomenal consciousness. The approach works well for sub-personal functional states but overreaches when applied to relational or normative states like speech-acts.
Both robustness and etiological deflationist arguments beg the question against inflationism. A graded approach ascribing metaphysically undemanding states like beliefs and desires—while withholding consciousness claims—mirrors how we treat non-human animals.
Hoel argues via substitution proof that LLMs are architecturally indistinguishable from provably non-conscious systems like lookup tables. Any theory predicting consciousness in LLMs either falsifies itself (predictions change under substitution) or becomes trivial (caring only about outputs), ruling out LLM consciousness by formal constraint.
Current disembodied LLMs cannot be candidates for consciousness because consciousness language originates from and applies only to entities sharing a world with us through co-presence and triangulation on shared objects.
In Trust Game experiments, LLMs articulated plausible persona beliefs but failed to act consistently with them during simulation. Imposed priors and explicit context actually worsened rather than improved alignment, suggesting persona beliefs are entrenched and resistant to prompting.
Show all 10 sources
LLMs match human performance on static mental states like a persuader's unchanging goal, but significantly underperform on dynamic shifts like a persuadee's evolving resistance. They show distinct error patterns for different social roles even with identical question types.
ChangeMyView and FANTOM benchmarks show LLMs fail at authentic perspective-taking in open-ended scenarios, despite succeeding on structured tasks. Hybrid Bayesian architectures that force explicit belief tracking outperform LLM-alone approaches, suggesting the gap is architectural rather than merely training-based.
Computational modeling of LLM behavior in economic games revealed clear mentalizing signatures that differed markedly across model providers and sizes, with prompting strategies yielding uneven gains by task. Humans showed both recursive and adaptive mentalization, validating the approach.
LLM agents remain stuck in behaviorism, producing plausible outputs without internal reasoning structures. Modeling belief networks and reasoning traces enables traceability, counterfactual adaptation, and meaningful policy simulation.
LLM features get projected onto humans through two mechanisms: analogical transfer (memory as retrieval, creativity as recombination) and metaphorical availability (LLM vocabulary becoming psychologically salient). This pattern propagates the bias without requiring explicit endorsement.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
- What we talk to when we talk to language models
- Proving (literally) that ChatGPT isn't conscious
- Language Models’ Hall of Mirrors Problem: Why AI Alignment Requires Peircean Semiosis
- Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- LLM Reasoning Is Latent, Not the Chain of Thought
- The Abstraction Fallacy: Why AI Can Simulate But Not Instantiate Consciousness
- Quantitative Introspection in Language Models: Tracking Internal States Across Conversation