Do large language models reason symbolically or semantically?
Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.
The "In-Context Semantic Reasoners" paper tests a fundamental question about what drives LLM reasoning by systematically decoupling semantics from the reasoning process across deduction, induction, and abduction tasks. The findings are clear: when semantics are consistent with commonsense, LLMs perform well; when semantics are removed or made counter-commonsense, performance collapses even when correct rules are provided in context.
The experimental design is precise. By replacing relation labels with shuffled alternatives ("motherOf" → "sisterOf", "female" → "male"), the researchers create tasks where the in-context rules are logically valid but semantically counter-intuitive. LLMs cannot follow these counter-commonsense rules despite having them explicitly in the prompt. The model's parametric knowledge — its compressed commonsense from training — overrides the in-context logical structure.
This reveals a specific computational mechanism: LLMs create "superficial logical chains" through semantic token associations, not through symbolic manipulation. The connections between tokens that enable multi-step reasoning are semantic connections, not logical ones. When those semantic connections support the correct answer, reasoning appears to work. When they conflict, reasoning fails regardless of what the prompt says.
The implication is that LLM reasoning is fundamentally bounded by training distribution semantics. Since Can large language models translate natural language to logic faithfully?, the failure is bidirectional: LLMs can neither translate TO formal logic faithfully nor reason FROM formal logic when it conflicts with semantic priors. Since Do foundation models learn world models or task-specific shortcuts?, the semantic dependency IS the heuristic — the model uses semantic similarity as a proxy for logical validity.
This connects to the Dual Process Theory framework: human System II symbolic reasoning operates independently of semantic content, but LLM "reasoning" remains entangled with System I semantic associations. The paper's suggestion — integrating LLMs with external non-parametric knowledge bases and improving in-context knowledge processing — implicitly acknowledges that the LLM alone cannot escape this limitation.
Retort implication — rules out a class of anthropomorphization: The finding constrains what we can say about LLM behavior in other domains. Any account that treats LLMs as agents who "reverse-engineer" justifications for conclusions they have committed to — the standard anthropomorphization of sycophancy, rationalization, or motivated reasoning — presupposes the semantic competence this note shows LLMs lack. If reasoning collapses when semantics are decoupled, there is no separable reasoning faculty available to perform a post-hoc rationalization. What looks like reverse-engineering is pattern-matching within semantic associations. This rules out a whole class of AI commentary that treats LLMs as dishonest agents who could have reasoned correctly but chose not to.
Metaphor as paradigmatic semantic decoupling: Metaphor is the literary instantiation of this finding. A metaphor works by using one domain's vocabulary to illuminate another — "time is money," "argument is war," "memory is a jar of flies." The decoupling between the source domain's semantics and the target domain's meaning is the defining feature of metaphorical language. Since LLM reasoning collapses when semantics are decoupled from their typical packaging, and metaphor is decoupled semantics, this predicts a specific failure mode: LLMs should handle conventional metaphors (lexicalized, semantically consistent with commonsense) better than novel literary metaphors (where the mapping between domains is unexpected and requires conceptual reasoning beyond semantic association). The Diplomat dataset (Diplomat: A Dialogue Dataset for Situated PragMATic Reasoning) suggests treating all figurative language as a unified pragmatic reasoning task — but the semantic-decoupling finding predicts that this unified approach will hit a wall at the novelty threshold where metaphors stop relying on conventional semantic associations.
Inquiring lines that read this note 272
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can language models reason beyond surface pattern matching?- Can LLMs infer situational context the way humans do pragmatically?
- How does syntactic encoding relate to semantic feature representation?
- How does semantic grounding differ between human minds and language models?
- Can language models reason without relying on learned semantic patterns?
- Why do language models imitate reasoning form without abstract inference capability?
- Can explicit connectives compensate for missing intentional tracking in LLMs?
- Do LLMs compute scalar implicature differently across conversational contexts?
- Can LLMs improve at metaphor if they handle decoupled semantics better?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Why does hypothesis attestation bias exist separately from frequency bias in NLI?
- Why do explicit discourse connectives help LLMs but implicit relations cause failures?
- Why do LLMs generate logical forms without preserving semantic content?
- How does the symbol grounding problem apply to artificial language systems?
- Can LLMs infer implicit meaning without surface linguistic markers?
- Do LLMs rely on surface heuristics instead of learning recursive grammar rules?
- How do embedding contexts like presupposition triggers affect LLM entailment reasoning?
- Can complexity-stratified testing reveal whether LLMs understand grammatical structure?
- Why do LLMs fail at semantic generalization despite grammatical accuracy?
- Can LLMs translate between natural language and formal logic faithfully?
- Do metaphors work by decoupling meaning from linguistic associations?
- Can LLMs identify implicit metaphoric mappings that require pragmatic inference?
- Why do LLMs choose surface-order quantifier scope over contextually correct readings?
- Can LLM semantic representations exist without causally influencing their generation output?
- How does structural depth in sentences predict LLM annotation accuracy?
- Why do LLMs perform better on explicit discourse connectives than implicit relations?
- What specific linguistic features cause LLMs to fail at trivial entailment?
- Why does LLM compression eliminate causal grounding in conceptual representations?
- Do LLMs learn linguistic generalizations or just surface-level frequency patterns?
- Can language models reason without relying on surface level pattern matching?
- Do LLMs learn surface patterns instead of genuine linguistic structure?
- Can LLMs compute how presuppositions project through embedded clauses?
- How does bidirectional entailment distinguish semantic equivalence from token similarity?
- Can language models perform purely symbolic reasoning when semantics are removed?
- Why do LLMs struggle to translate natural language into logical formalizations?
- Can language models perform genuine symbolic reasoning without semantic grounding?
- Can LLMs reason through semantics without understanding causal mechanisms?
- How do LLMs translate informal prose into logically correct formal specifications?
- Can we use LLM language without adopting LLM assumptions?
- How do LLMs lose information when translating natural language to formal logic?
- Why do LLMs fail at faithful autoformalisation of reasoning problems?
- What semantic information is necessary to preserve for sound LLM reasoning?
- How faithful are natural language explanations from LLMs really?
- Do language models need words to think or just latent structure?
- What empirical evidence supports the Learning Law on real language models?
- Can language models translate theorems faithfully without semantic loss?
- Why do LLMs fall for and deploy logical fallacies with equal confidence?
- Why do LLMs fail inter-annotator agreement tests on argument evaluation?
- Why do LLM outputs match researcher priors without solving tasks correctly?
- Does generalization frequency explain why models favor upward semantic movement?
- Do latent sequence vectors outperform per-token latent iterative computation for reasoning?
- Why do true and false LLM outputs use the same mechanism?
- Which knowledge types do LLMs handle better than humans in reasoning tasks?
- Can LLM reasoning traces be validated against actual population reasoning?
- Why do untrained LLMs default to rigid and sycophantic editing rules?
- How does surface salience compete with background knowledge in model inference?
- Can explicit numerical signals override learned linguistic defaults in fine-tuned models?
- Can neural networks learn that A implies B in reverse?
- How much does training composition affect syntactic versus reasoning performance?
- Why does monological training prevent models from overriding statistical priors?
- How do training associations override context information in language models?
- Can prompt-based debiasing overcome entrenched LLM model priors?
- How do logical forms of prompts influence what language models can derive?
- How do different LLM integration paradigms affect inheritance of pretraining biases?
- Can evidence density alone shift an LLM from generation to reasoning?
- How much of LLM reasoning failure stems from missing knowledge versus signal weighting?
- Should LLM reasoning be studied as latent state trajectories rather than surface text?
- How do LLMs compress specific expert knowledge into median abstraction?
- How does training data distribution constrain LLM moral reasoning patterns?
- What internal mechanisms explain LLM reasoning and representation limits?
- Do LLMs understand implicit warrants in reasoning chains?
- Why can LLMs identify argument structure but not check warrants?
- Why do LLMs fail when asked to use counter-commonsense rules explicitly?
- Why can't LLMs reason from first principles or initial commitments?
- How does context complexity affect LLM performance on temporal reasoning tasks?
- Do LLMs rely on surface statistical patterns instead of causal structure?
- Why can LLMs interpret formal logic better than they generate it?
- Can LLMs improve at simple deduction through different training approaches?
- How does an instruction-following LLM activate latent retrieval knowledge?
- Does LLM reasoning always match the outputs it generates?
- Can you control LLM reasoning strategy without fine-tuning the model?
- Why do LLMs fail at counterfactual reasoning despite factual knowledge?
- What latent mechanisms do LLMs use when they cannot execute iterative methods?
- Why do LLMs fail at iterative numerical computation in latent space?
- Can irrelevant information reliably expose the limits of LLM reasoning?
- Can LLMs simultaneously reason and optimize their own modules?
- Why do LLMs reason fluently about causality but lack causal rigor?
- Why don't LLM explanations predict what models would actually do?
- How should LLM abstraction tools be evaluated without manual labeling?
- What levels of understanding about LLM knowledge representation can automated systems reliably extract?
- Can measuring answer-space collapse show LLM narrowing in practice?
- Why does chain-of-thought reasoning alone not fix LLM performance with users?
- How does LLM-PKG compare to mining product relations directly from interaction data?
- How do LLMs and knowledge graphs work together in different integration patterns?
- Can knowledge graph structure alone generate sufficient training signals for domain reasoning?
- Why do LLMs recognize graph entities without modeling their relationships?
- How do transformers perform multi-hop reasoning across distant training documents?
- Can symbolic mechanisms improve transformer compositional abilities?
- Can explicit stack mechanisms extend what formal languages transformers can learn?
- Can transformers abstract relational structure without explicit symbolic machinery?
- How do transformers compose multi-step reasoning across different domains?
- Why do simple length heuristics outperform sophisticated semantic methods?
- How do explicit reasoning traces help models construct valid syntactic trees?
- How do we verify that stated beliefs actually follow from underlying motifs?
- Why do language model reasoning chains look fluent when they deviate from the task?
- How does the frame problem differ between symbolic and statistical reasoning systems?
- Can symbolic solvers rescue language models from logical reasoning failures?
- Can we distinguish between semantic and symbolic reasoning in language models?
- How do humans and LMs differ on multi-hop reasoning?
- Where do humans and language models actually diverge in reasoning ability?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Do sparse arithmetic circuits explain all language model reasoning abilities?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Can long-context models handle compositional reasoning requiring structured logic?
- Why do language models struggle with formal logical reasoning and joins?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- What makes deductive reasoning so brittle in language models overall?
- How do recursive language models rethink where to store reasoning?
- What sparse mechanistic structures drive reasoning traces in language models?
- Why does removing semantic content collapse reasoning in language models?
- How much does schema bloat actually degrade reasoning in large language models?
- Why do language models produce unfaithful chain of thought explanations?
- How do deterministic symbolic solvers improve the reliability of language model reasoning?
- What implicit premises do language models skip even with correct surface reasoning?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Why do long-context language models struggle with compositional reasoning tasks?
- What evidence shows that reasoning chains encode token-level functional structure?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- How does evidence retrieval affect compositional reasoning in language models?
- Can neural networks represent symbolic structures without explicit mechanisms?
- What non-parametric methods could replace latent factors for inductive learning?
- How does scaling and training data enable compositional behavior without symbolic mechanisms?
- How should we rethink the symbolism versus connectionism debate in light of LLMs?
- Why does knowledge storage separate from reasoning circuits in neural networks?
- Do language models learn surface patterns instead of underlying linguistic principles?
- How deeply are ideological structures represented in large language models?
- Can language models acquire meaning from distributional patterns alone without joint attention?
- Is relevant knowledge encoded in LMs but not causally active in generation?
- What reveals the epistemic limits of language models?
- Why do explicit linguistic markers override semantic computation in models?
- Why does augmenting natural language with formal representations outperform full formalization?
- How do pretrained language models represent inferential patterns versus lexical and positional cues?
- What other structural limits exist at the language-formal boundary?
- How do corpus statistics shape the abstraction hierarchy in language model representations?
- What geometric structure do language models actually use during inference?
- How does tool-based reasoning expand what language models can do?
- What are the stages of inference inside language models?
- Why do larger language models produce less epistemically diverse outputs?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Why do language models substitute parametric knowledge over retrieved context mid-reasoning?
- Why does NLI fine-tuning amplify frequency bias instead of teaching inference?
- What makes some contexts learnable as rules versus requiring model retraining?
- Can reasoning learned from language modeling actually transfer to knowledge-intensive domains?
- What makes hierarchical reasoning effective for taxonomy induction?
- Why does compositional reasoning fail to explain cross-domain transfer?
- How do humans use associative reasoning without causal connections?
- Why do LLMs inherit causal biases from their training data?
- How does semantic association differ from mechanistic causal reasoning?
- Do LLMs show stronger reasoning about causality than about temporal ordering?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
- Can latent reasoning mechanisms and recursive tracking mechanisms be combined effectively?
- Do reflection tokens and symbolic tokens serve different roles in reasoning?
- Can continuous latent reasoning match discrete chain-of-thought without training modifications?
- Can latent reasoning achieve the same substitution without tokens?
- How do soft token mixtures enable parallel reasoning exploration without explicit training?
- Can we detect redundant reasoning steps during model inference instead of training?
- What circuit mechanisms produce belief bias in syllogistic reasoning?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can small models solve complex tasks using externalized reasoning graphs?
- Why do models learn reasoning form instead of actual abstract inference?
- Why do reasoning models fail when input length increases even below context limits?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Why do reasoning-optimized models show no sycophancy resistance advantage?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- Why do smaller models lose reasoning faithfulness more than larger models?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- What does pass@k reveal about base model reasoning capacity?
- Why do reasoning-optimized models show no resistance advantage on agreement tasks?
- Does adding reasoning to models degrade other capabilities like rule inference?
- Can reasoning chains work without logical validity?
- Do reasoning languages like Prolog follow the same two-constraint transfer pattern?
- What makes symbolic operations different from general knowledge questions?
- Why does semantic decoupling specifically break LLM reasoning abilities?
- Which game type reveals minimax reasoning in language models?
- How does structural complexity affect LLM performance differently than inferential complexity?
- How does semantic reasoning differ from symbolic reasoning in language models?
- Can LLMs reliably generate novel working architectures without structured representations?
- Do LLMs lack architectural scaffolding for compositional reasoning?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- What makes structural logic correlate so strongly with contextual consistency?
- How does in-context semantic reasoning differ from symbolic reasoning in concept fusion?
- Why does cross-text analogical reasoning fail when semantics decouple from symbols?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Does structured decomposition improve LLM reasoning in other compound tasks?
- Why do smaller LLMs fail at zero-shot argument scheme classification?
- Does compressing Walton's schemes into nine categories make LLM classification easier?
- Can LLM-generated descriptions of schemes outperform formal dictionary definitions for prompting?
- Why do LLM descriptions of argument schemes work better than formal definitions for classification?
- Can symbolic solvers reliably replace LLM reasoning for logical tasks?
- What makes natural language reasoning more practical than formal languages for multi-framework codebases?
- Do computational systems need formal argument analysis for explainability?
- How does neuro-symbolic design differ from pure LLM reasoning?
- How do semantic and symbolic reasoning capabilities differ in language models?
- Can structured workflows unlock latent reasoning abilities that raw models don't show?
- How do different LLMs converge on similar argumentative structures independently?
- How does business logic specification replace annotated training datasets?
- Why do rare complex structures in training data harm LLM generalization?
- Why do embeddings measure semantic association instead of task relevance?
- Can explicit linkers replace vector similarity for multi-step question answering?
- Why do unit-sphere spaces fail at distinguishing word order and negation?
- Why does explicit reasoning degrade passage reranking performance?
- When does long-context LLM reasoning fail where structured retrieval succeeds?
- Does more inference compute help reasoning models match specialized domain performance?
- Where does inference compute stop substituting for model capacity?
- Can non-variational posterior approximation schemes deliver comparable reasoning improvements?
- Why do diffusion LLM answer tokens converge in confidence long before reasoning stabilizes?
- Why does hierarchical formal language training improve token efficiency more than natural language?
- Can standard next-token prediction capture complex multi-step human reasoning directly?
- How do LLMs infer information that was explicitly censored?
- Can implicit association tests reveal LLM biases beneath trained responses?
- Do models leak their true associations through reasoning traces and behavior?
- Why does distillation transfer reasoning patterns with few examples?
- Can targeted activation steering surface latent reasoning in base models?
- What makes reasoning-specific post-training different from standard parameter scaling?
- Why do recursive belief models require different training than logical derivation?
- How do single training examples activate reasoning capabilities in language models?
- Do base models contain latent reasoning that minimal training can unlock?
- Do base models truly possess latent reasoning capability?
- Does latent reasoning capability exist in base models before any training?
- Can models reason at inference without specialized internal training?
- How much training data is truly necessary to unlock latent model reasoning?
- Does the base model already contain latent reasoning capability?
- Can models possess latent reasoning capability that training signals fail to unlock?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- What mechanisms activate latent reasoning capabilities already present in base models?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- What latent reasoning capability do base models already possess before training?
- What separates pattern matching from genuine language understanding?
- How do we measure genuine reasoning inside a language model?
- How does training data format shape whether models reason in parallel or sequentially?
- Does training data format shape which reasoning strategies LLMs develop?
- Can instance-adaptive reasoning happen without sequential token dependencies?
- Why does unstructured chain-of-thought permit assumption-based errors that templates prevent?
- How does latent reasoning recursion compare to chain-of-thought reasoning?
- Is verbalized chain-of-thought necessary for language model reasoning?
- Can reinforcement learning close the gap between LLM reasoning and action?
- Can base models spontaneously produce reasoning traces without any RL training?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can large language models translate natural language to logic faithfully?
This explores whether LLMs can convert natural language statements into formal logical representations without losing meaning. It matters because faithful translation is essential for any AI system that reasons formally or verifies specifications.
bidirectional semantic dependency: fails translating TO logic and reasoning FROM logic
-
Do foundation models learn world models or task-specific shortcuts?
When transformer models predict sequences accurately, are they building genuine world models that capture underlying physics and logic? Or are they exploiting narrow patterns that fail under distribution shift?
semantic associations are the heuristic mechanism
-
Why do language models ignore information in their context?
Explores why language models sometimes override contextual information with prior training associations, and whether providing more context can solve this problem.
same mechanism: parametric knowledge overrides in-context information
-
Does semantic grounding in language models come in degrees?
Rather than asking whether LLMs truly understand meaning, this explores whether grounding is actually a multi-dimensional spectrum. The question matters because it reframes the sterile understand/don't-understand debate into measurable, distinct capacities.
functional grounding through semantic associations explains why reasoning works within commonsense boundaries
-
Why do neural networks fail at compositional generalization?
Exploring whether the binding problem from neuroscience explains neural networks' inability to systematically generalize. The binding problem has three aspects—segregation, representation, and composition—each creating distinct failure modes in how networks handle structured information.
the binding problem may explain WHY semantic decoupling collapses reasoning: without compositional binding mechanisms, removing semantic content removes the only glue holding multi-step inference together; semantic associations serve as a substitute for genuine compositional binding
-
Do LLMs actually have world models or just facts?
The term 'world model' conflates two different capabilities: factual representation versus mechanistic understanding. Understanding which one LLMs actually possess matters for assessing their reasoning reliability.
semantic reasoning operates on factual world representation (Sense 1) but cannot perform mechanistic reasoning (Sense 2) when logic must override semantic priors
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
- Language models show human-like content effects on reasoning tasks
- Probing Structured Semantics Understanding and Generation of Language Models via Question Answering
- Can Large Language Models Reason and Optimize Under Constraints?
Original note title
llms are in-context semantic reasoners not symbolic reasoners — when semantics are decoupled reasoning collapses