SYNTHESIS NOTE
Topics›Reasoning Critiques›this note

Does chain-of-thought reasoning actually generalize beyond training data?

Explores whether CoT's strong performance on benchmarks reflects genuine reasoning ability or merely reflects learned patterns tied to specific distributions. Tests how CoT behaves when tasks, formats, or reasoning length shift away from training data.

Synthesis note · 2026-02-22 · sourced from Reasoning Critiques

Chain-of-Thought prompting performs well on in-distribution problems and fails predictably as distributional discrepancy increases. This is not a bug — it is the fundamental nature of what CoT is.

DataAlchemy experiments train LLMs from scratch in controlled environments and probe them under three distributional shift dimensions:

  1. Task distribution shift — novel tasks with unique elements or underlying logical structure not seen during training
  2. Length distribution shift — reasoning chains substantially longer or shorter than training data length range
  3. Format distribution shift — prompt formulation variations (even minor syntactic changes) that fall outside training distribution

In all three dimensions, the pattern is the same: CoT works within distribution, fails outside it. Under moderate shifts, models generate fluent yet logically inconsistent reasoning — the form holds, the logic breaks. This is the "mirage" phenomenon: outputs look like reasoning while producing wrong conclusions.

The interpretive frame: CoT reflects a structured inductive bias learned from training data, not a generalizable reasoning capability. When a test query is within this inductive bias, CoT activates the appropriate reasoning schema and produces good outputs. When the query falls outside it, the schema mismatch produces confident-sounding nonsense.

The practical implication for CoT as a plug-and-play solution: it is not. Performance on CoT benchmarks measures in-distribution capability. Extrapolating to novel tasks, unusual prompt formulations, or unusually long/short reasoning chains is unjustified. The benchmark scores do not predict performance under distribution shift.

This provides the empirical grounding for Does chain-of-thought reasoning reveal genuine inference or pattern matching? — the mirage emerges from imitation under distribution shift: the model continues imitating the form of reasoning while having no schema to produce valid content.

Inquiring lines that read this note 276

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What prevents LLMs from applying their reasoning knowledge to improve outputs? When do simpler collaborative filtering approaches outperform complex LLM recommenders? What explains the gap between benchmark scores and true reasoning capability? What prevents language models from performing systematic logical reasoning? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Can latent reasoning match or exceed explicit reasoning performance? Can AI systems evade safety evaluations through reasoning manipulation? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? How do thinking tokens exhibit diminishing returns in reasoning? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Can mechanistic interpretability methods reliably reveal what models actually know? When should retrieval systems decide to fetch new information? Can minimal training unlock latent reasoning already present in base models? Does augmenting symbolic reasoning improve LLM logical reasoning ability? How does fine-tuning trade off accuracy against reasoning quality? How does model capacity affect learning performance on diverse downstream tasks? Can inference-time computation adaptively substitute for static model capacity? How reliably can language models perform causal versus temporal reasoning? Does training data format shape model reasoning more than domain content? How do real-world evaluations reveal AI capabilities that benchmarks hide? Why do models reveal hidden associations despite concealment attempts? How do interpretive frames override surface features in text comprehension? What makes reasoning traces effective supervision even when they're incorrect? How do sequence length and task type interact with sparsity tolerance? What are the fundamental limits of prompting for language models? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Can reasoning traces reveal actual model reasoning versus plausible output? What limits language model accuracy in evaluating ideas? When does parallel reasoning outperform sequential reasoning with the same token budget? How do training data quality and composition affect downstream model performance? What gaps exist between benchmark performance and real deployment outcomes? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Why do training associations persist despite contradictory contextual information? Should models ask for clarification when facing ambiguous or under-specified information? How do transformer attention patterns implement retrieval and reasoning? What makes process supervision effective for training complex reasoning models? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Why don't better reasoning capabilities improve theory of mind performance? Why do multi-agent systems reach premature consensus without genuine deliberation? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Can AI agents improve their skills through accumulated experience and reuse? How do agents learn to distinguish valuable feedback from noise? Do accumulated memories help or hurt continual learning in models? How do curriculum design and feedback approaches affect model learning? What prediction granularity best trains models to generate reliable reasoning? How do neural networks learn compositional structure from training? Should agents compress episodic memory or retain raw interaction histories? Why do abstract preferences outperform episodic memories in personalization? Can base models hide emergent misalignment through alignment training? What human oversight must AI research systems have? How does optimization for reward create emergent misalignment in language models? How do users confuse explanation quality with actual system accuracy? How does decomposing tasks into separate stages affect reasoning quality and safety? Can external verification systems adequately replace learned reasoning in AI outputs? How should retrieval strategies adapt to multi-step reasoning demands? How can AI systems reliably guide voters without introducing political bias? Can AI systems achieve real improvement without external human feedback? How do AI systems determine and balance multiple competing objectives?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cot reasoning is distribution-bounded — effectiveness degrades predictably with distributional discrepancy