SYNTHESIS NOTE
Topics›Reasoning Logic Internal Rules›this note

Do large language models reason symbolically or semantically?

Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.

Synthesis note · 2026-02-22 · sourced from Reasoning Logic Internal Rules

The "In-Context Semantic Reasoners" paper tests a fundamental question about what drives LLM reasoning by systematically decoupling semantics from the reasoning process across deduction, induction, and abduction tasks. The findings are clear: when semantics are consistent with commonsense, LLMs perform well; when semantics are removed or made counter-commonsense, performance collapses even when correct rules are provided in context.

The experimental design is precise. By replacing relation labels with shuffled alternatives ("motherOf" → "sisterOf", "female" → "male"), the researchers create tasks where the in-context rules are logically valid but semantically counter-intuitive. LLMs cannot follow these counter-commonsense rules despite having them explicitly in the prompt. The model's parametric knowledge — its compressed commonsense from training — overrides the in-context logical structure.

This reveals a specific computational mechanism: LLMs create "superficial logical chains" through semantic token associations, not through symbolic manipulation. The connections between tokens that enable multi-step reasoning are semantic connections, not logical ones. When those semantic connections support the correct answer, reasoning appears to work. When they conflict, reasoning fails regardless of what the prompt says.

The implication is that LLM reasoning is fundamentally bounded by training distribution semantics. Since Can large language models translate natural language to logic faithfully?, the failure is bidirectional: LLMs can neither translate TO formal logic faithfully nor reason FROM formal logic when it conflicts with semantic priors. Since Do foundation models learn world models or task-specific shortcuts?, the semantic dependency IS the heuristic — the model uses semantic similarity as a proxy for logical validity.

This connects to the Dual Process Theory framework: human System II symbolic reasoning operates independently of semantic content, but LLM "reasoning" remains entangled with System I semantic associations. The paper's suggestion — integrating LLMs with external non-parametric knowledge bases and improving in-context knowledge processing — implicitly acknowledges that the LLM alone cannot escape this limitation.

Retort implication — rules out a class of anthropomorphization: The finding constrains what we can say about LLM behavior in other domains. Any account that treats LLMs as agents who "reverse-engineer" justifications for conclusions they have committed to — the standard anthropomorphization of sycophancy, rationalization, or motivated reasoning — presupposes the semantic competence this note shows LLMs lack. If reasoning collapses when semantics are decoupled, there is no separable reasoning faculty available to perform a post-hoc rationalization. What looks like reverse-engineering is pattern-matching within semantic associations. This rules out a whole class of AI commentary that treats LLMs as dishonest agents who could have reasoned correctly but chose not to.

Metaphor as paradigmatic semantic decoupling: Metaphor is the literary instantiation of this finding. A metaphor works by using one domain's vocabulary to illuminate another — "time is money," "argument is war," "memory is a jar of flies." The decoupling between the source domain's semantics and the target domain's meaning is the defining feature of metaphorical language. Since LLM reasoning collapses when semantics are decoupled from their typical packaging, and metaphor is decoupled semantics, this predicts a specific failure mode: LLMs should handle conventional metaphors (lexicalized, semantically consistent with commonsense) better than novel literary metaphors (where the mapping between domains is unexpected and requires conceptual reasoning beyond semantic association). The Diplomat dataset (Diplomat: A Dialogue Dataset for Situated PragMATic Reasoning) suggests treating all figurative language as a unified pragmatic reasoning task — but the semantic-decoupling finding predicts that this unified approach will hit a wall at the novelty threshold where metaphors stop relying on conventional semantic associations.

Inquiring lines that read this note 272

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can language models reason beyond surface pattern matching? What limits language model accuracy in evaluating ideas? Why do training associations persist despite contradictory contextual information? What are the fundamental limits of prompting for language models? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Why does polished AI output gain credibility despite fundamental verifiability problems? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? How do transformer attention patterns implement retrieval and reasoning? Can reasoning traces reveal actual model reasoning versus plausible output? What prevents language models from performing systematic logical reasoning? How do neural networks learn compositional structure from training? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How does fine-tuning trade off accuracy against reasoning quality? Can AI systems discover fundamental improvements to their own architectures? How reliably can language models perform causal versus temporal reasoning? Can latent reasoning match or exceed explicit reasoning performance? What capabilities differentiate diffusion from autoregressive language models? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Does augmenting symbolic reasoning improve LLM logical reasoning ability? How do training data quality and composition affect downstream model performance? Why do vector embeddings fail at capturing task-relevant relationships? How does diversity prevent model convergence on superficial patterns? How should retrieval strategies adapt to multi-step reasoning demands? Can inference-time computation adaptively substitute for static model capacity? What prediction granularity best trains models to generate reliable reasoning? Why do models reveal hidden associations despite concealment attempts? Do language models reason through disagreement or only accommodate it? Can minimal training unlock latent reasoning already present in base models? Should models ask for clarification when facing ambiguous or under-specified information? Does training data format shape model reasoning more than domain content? Can humans reliably detect and resist AI-generated misinformation? Can AI systems achieve real improvement without external human feedback? What structural biases does transformer attention architecture inherently introduce? Why do language models hallucinate and how can we prevent it? Does intelligent routing among smaller models outperform training larger models? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Can language models reliably simulate personas and predict behavior? How does scaling reasoning capabilities affect models' appropriate abstention behavior? What causes coordination failures in multi-agent language model systems? What distinguishes genuine communicative competence from surface language performance? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Why do LLM research ideation systems generate novelty but lack diversity? Can models develop genuine introspective capability, or only mimic it? Can mechanistic interpretability methods reliably reveal what models actually know? How can we reduce inherent biases in LLM-based evaluation judges? Can we trust AI-generated mathematical proofs without understanding them? Can external verification systems adequately replace learned reasoning in AI outputs?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
22 direct connections · 191 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llms are in-context semantic reasoners not symbolic reasoners — when semantics are decoupled reasoning collapses