Does chain-of-thought reasoning actually generalize beyond training data?
Explores whether CoT's strong performance on benchmarks reflects genuine reasoning ability or merely reflects learned patterns tied to specific distributions. Tests how CoT behaves when tasks, formats, or reasoning length shift away from training data.
Chain-of-Thought prompting performs well on in-distribution problems and fails predictably as distributional discrepancy increases. This is not a bug — it is the fundamental nature of what CoT is.
DataAlchemy experiments train LLMs from scratch in controlled environments and probe them under three distributional shift dimensions:
- Task distribution shift — novel tasks with unique elements or underlying logical structure not seen during training
- Length distribution shift — reasoning chains substantially longer or shorter than training data length range
- Format distribution shift — prompt formulation variations (even minor syntactic changes) that fall outside training distribution
In all three dimensions, the pattern is the same: CoT works within distribution, fails outside it. Under moderate shifts, models generate fluent yet logically inconsistent reasoning — the form holds, the logic breaks. This is the "mirage" phenomenon: outputs look like reasoning while producing wrong conclusions.
The interpretive frame: CoT reflects a structured inductive bias learned from training data, not a generalizable reasoning capability. When a test query is within this inductive bias, CoT activates the appropriate reasoning schema and produces good outputs. When the query falls outside it, the schema mismatch produces confident-sounding nonsense.
The practical implication for CoT as a plug-and-play solution: it is not. Performance on CoT benchmarks measures in-distribution capability. Extrapolating to novel tasks, unusual prompt formulations, or unusually long/short reasoning chains is unjustified. The benchmark scores do not predict performance under distribution shift.
This provides the empirical grounding for Does chain-of-thought reasoning reveal genuine inference or pattern matching? — the mirage emerges from imitation under distribution shift: the model continues imitating the form of reasoning while having no schema to produce valid content.
Inquiring lines that read this note 276
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What prevents LLMs from applying their reasoning knowledge to improve outputs?- What makes conceptual inquiry the fastest high-scoring AI interaction pattern?
- Can frozen world models from training cutoff remain adequate for real-world reasoning?
- Can verification loops and decomposition fix judgment failures?
- Why do naive baselines outperform trained models in entity-level CRS evaluation?
- Why do cross-product features fail to generalize across unseen feature combinations?
- What other hidden biases might aggregate metrics fail to distinguish from reasoning?
- How should we redesign benchmarks to catch conservative bias in reasoning tasks?
- Does the Heuristic Override Benchmark measure enumeration or world knowledge?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- How do humans and LMs differ on multi-hop reasoning?
- Can reasoning benchmarks separate logic from believability?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Why does comparison reasoning generalize better than composition reasoning?
- Which RAG sub-decisions are actually pattern matching versus reasoning intensive?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Why are pairwise relations insufficient for representing higher-order multi-hop reasoning?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Why does second-hop reasoning fail when composed with out-of-distribution triples?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- How does instance novelty rather than chain length explain reasoning failure?
- Does chain-of-thought text causally drive reasoning or merely reflect it?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Is chain-of-thought reasoning actual computation or distribution imitation?
- What happens to chain-of-thought performance across distribution shifts?
- How do chain-of-thought structures affect reasoning robustness?
- How does chain-of-thought training change higher layer computations?
- Can chain of thought reasoning actually validate logical arguments?
- Does chain-of-thought reasoning specifically improve performance on metalinguistic tasks?
- Does chain-of-thought reasoning improve mental state tracking in dialogue?
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- How does graph of thoughts enable divide-and-conquer reasoning patterns?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- How does chain-of-thought reasoning become decorative after domain-specific fine-tuning?
- Does chain-of-thought reasoning help or hurt social reasoning tasks?
- Why does extended chain-of-thought reasoning fail to improve numerical optimization performance?
- Why does chain-of-thought fail to improve multimodal model perception performance?
- What distinguishes graph-of-thought reasoning from other structured reasoning topologies?
- What makes token-level reasoning during pretraining different from test-time chain-of-thought?
- Does chain-of-thought accuracy degrade with longer reasoning traces?
- Does CoT reasoning actually cause the outputs that follow it?
- Can single representation edits match chain-of-thought reasoning without explicit steps?
- How does trajectory geometry relate to the need for chain-of-thought reasoning?
- Do gold CoT tokens avoid the need for specialized training data?
- When does multi-hop reasoning improve chain-of-thought monitor detection?
- Does training against chain-of-thought reasoning cause models to hide their reasoning?
- How much does chain-of-thought reasoning actually determine model outputs?
- Can steering a single latent feature replicate chain-of-thought performance?
- Can step-level deliberation flags guide other reasoning systems?
- Does iterative denoising order affect the reasoning style diffusion models learn?
- Does changing decoding procedure reveal hidden chain-of-thought paths?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- Can extended thinking genuinely improve reasoning or just increase variance?
- How do gradient descent iterations at inference compare to chain-of-thought reasoning chains?
- When does explicit reasoning actually degrade performance on a task?
- Can latent reasoning in continuous space scale beyond supervised reasoning tasks?
- Can extended reasoning training capture individual strategic thinking styles?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can we transfer reasoning structure without copying surface form?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- Can recursive subtask trees implement tree-of-thought reasoning more efficiently?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- Can dataset design systematically expand reasoning graph diameter?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Can continuous latent reasoning match discrete chain-of-thought without training modifications?
- How much reasoning depth do we actually need for most real-world tasks?
- Can we improve reasoning by amplifying information at mutual information peaks?
- Can memorization scores diagnose where reasoning chains become unreliable?
- Can minimal reasoning steps match verbose reasoning accuracy?
- Does reasoning style transfer matter more than solution correctness in distillation?
- What computational structures can actually scale serial reasoning depth?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Why does distilling reasoning strategies outperform raw trajectory memory?
- Does the latent-explicit gap widen beyond 3B parameters on reasoning tasks?
- What detection methods can catch each distinct CoT bypass strategy?
- Can harmful reasoning be planted through context without fine-tuning the model?
- How do transformers perform multi-hop reasoning across distant training documents?
- How sensitive is analogical reasoning emergence to training data and scale?
- Why does scaling reasoning tokens fail to improve unfamiliar tasks?
- Why does extended thinking increase output variance without improving reasoning quality?
- What makes a background condition relevant to a specific reasoning task?
- Why do models fail on logically equivalent tasks with different data distributions?
- Can small models solve complex tasks using externalized reasoning graphs?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Does reasoning structure match explicit versus implicit task demands?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Is the reasoning cliff actually a tool-use problem?
- How does scaling reasoning capability actually reduce instruction-following ability?
- Can benchmark improvements hide degradation of deliberative reasoning?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- What does pass@k reveal about base model reasoning capacity?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Why does strategy diversity within reasoning chains improve model generalization?
- How does task simplification affect analogical reasoning patterns?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- What distinguishes task-specific heuristics from genuine world models?
- What role does embedding space geometry play in multi-hop reasoning?
- What makes reasoning capability a pre-training rather than post-training phenomenon?
- How much does pre-training frequency predict reasoning task performance?
- How much does pretraining contribute to ToM performance versus task-specific training?
- Can reasoning skills trained on law improve performance in STEM?
- Why does distillation transfer reasoning patterns with few examples?
- What makes reasoning-specific post-training different from standard parameter scaling?
- Can models trained on longer contexts develop better fundamental reasoning abilities?
- How does a single training example trigger phase transitions in reasoning output?
- Does this reasoning steering method work consistently across all model sizes?
- Why do instruction following and reasoning capability trade off in training?
- How can one training example improve reasoning across thousands of unseen problems?
- Does penalizing thought transitions improve reasoning without model retraining?
- What makes thought identifiability provable without auxiliary training data?
- Why does reasoning training improve math but hurt knowledge tasks?
- Can activation steering vectors compress reasoning without retraining models?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- Does latent reasoning capability exist in base models before any training?
- How do timing and search internalization interact during reasoning post-training?
- Why does reasoning transfer across different numbers but factual recall does not?
- Can smaller amounts of diverse reasoning demonstrations replace exhaustive factual training data?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Does token-level reasoning during pretraining improve general reasoning without task-specific supervision?
- How does RPT compare to learning when versus how to deploy reasoning?
- Can distillation from stronger models create genuinely new reasoning abilities?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- Why does single-shot learning fail in REVTHINK's multi-source reasoning tasks?
- Can articulating latent reasoning processes improve transfer across domains?
- Can small demonstration sets unlock general reasoning without large question data?
- Why do reasoning gains resist clear attribution to specific training changes?
- What makes some reasoning strategies genuinely novel versus latent?
- Does looped pretraining build reasoning more efficiently than supervised fine-tuning?
- What formal representation could capture analogical reasoning across domains?
- Do reasoning systems reuse cognitive structures across unrelated topics?
- How do game-based benchmarks reveal reasoning fragmentation across domains?
- What makes multi-paradigm chaining a distinct reasoning topology?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- What distinguishes the convergence patterns between reasoning and lexical variation tasks?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- Does domain training degrade reasoning ability even when benchmark scores rise?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How does cross-domain reasoning transfer differ from domain-specific knowledge transfer?
- What makes knowledge-rich specialized domains structurally different from general reasoning tasks?
- What makes certain bond distributions more learnable than others?
- Why does SFT reduce reasoning quality even when improving domain accuracy?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Can attribute decomposition improve other interactive reasoning tasks beyond clinical questioning?
- Why does semantic similarity retrieval enable skill transfer to novel situations?
- Can mathematical reasoning improvements transfer across problem subdomains?
- How does supervised fine-tuning degrade chain-of-thought faithfulness over time?
- Why do non-experts default to familiar chart types despite domain complexity?
- Does task diversity in pretraining data transfer reasoning better than larger models?
- What makes procedural knowledge in documents generalize better than facts?
- Can expert-derived knowledge bases scale to other high-stakes domains?
- How much of the combinatorial task space must training data cover?
- Does compositional generalization emerge suddenly or improve smoothly with scale?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- Can scaling data alone solve performance gaps on long-tail concepts?
- How does inference compute substitution affect the training parameter scaling trade-off?
- Can test-time scaling prioritize genuine reasoning over pattern matching?
- Can adaptive compute distribution across prompts replace the need for sophisticated reasoning frameworks?
- Does more inference compute help reasoning models match specialized domain performance?
- What mechanisms drive test-time compute allocation in reasoning tasks?
- How much does test-time compute improve reasoning without more tokens?
- Why do long-horizon reasoning tasks need per-turn step limits rather than just compute budgets?
- Does inference-time compute improve pretraining data efficiency in practice?
- What inference-time scaling benefits emerge from reasoning before each prediction?
- What patterns emerge across test-time scaling and reasoning architectures?
- Can reasoning models outperform non-reasoning models with more inference compute?
- Can activation patching reveal which reasoning steps actually matter?
- What makes counterfactual thinking different from behavioral pattern matching?
- What makes a causal abstraction more transferable than a generic heuristic?
- How much does training data format shape what reasoning strategy emerges?
- Why does training format shape reasoning strategy more than domain?
- Why does training data format shape reasoning strategy more than domain content?
- Does training data format shape model reasoning more than domain content?
- How does training format shape reasoning strategy more than content?
- How does training data format shape whether models reason in parallel or sequentially?
- How much does training data presentation format shape reasoning ability?
- How does training data format shape which reasoning patterns emerge in models?
- Why does training data format shape reasoning strategy more than content?
- Can training format itself shape what reasoning strategy a model learns?
- Does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- How does training data structure shape reasoning strategy more than domain content?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How should domain-specific AI be evaluated differently from general benchmarks?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How can high benchmark performance mask broken reasoning in AI systems?
- Do automated benchmarks accurately measure real-world strategic reasoning ability?
- Why does fine-tuning sometimes damage chain-of-thought reasoning even when accuracy improves?
- How does post-training on traces improve performance without semantic reasoning?
- Can deliberate corruption of reasoning traces harm out of distribution generalization?
- Why does long CoT training optimize for structural coherence over content correctness?
- Does trace length actually reflect problem difficulty or training proximity?
- Why do shorter confident reasoning traces fail on out-of-distribution problems?
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- Does reasoning training create blind spots in premise detection?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Why do task-specific heuristics fail at generalizing to sparse data regions?
- Does sparsity-guided ordering work equally well for reasoning and classification tasks?
- Does irrelevant context degrade reasoning even within model context limits?
- Which structural properties of CoT prompts matter most for performance?
- Can hyperedges replace triple-based externalization in reasoning tasks?
- Can knowledge graphs externalize and validate reasoning steps during inference?
- Does small-world structure in reasoning graphs improve generalization?
- Can knowledge graph structure alone generate sufficient training signals for domain reasoning?
- How do random walk reasoning chains from knowledge graphs compare to traditional fine-tuning?
- Can curriculum graphs as training data improve model understanding of prerequisite chains?
- What saliency patterns distinguish successful from failed chain-of-thought reasoning?
- Why do we measure reasoning quality by reading visible chains?
- What metric distinguishes deep reasoning from superficial information propagation?
- Do longer chain-of-thought traces improve interpretability or just performance?
- How much do compressed reasoning traces transfer across different models?
- Can post-hoc analysis of reasoning traces actively mislead users?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- Can breadth-first search in continuous space outperform chain-of-thought on logical tasks?
- How does meta-reasoning combine information distributed across multiple chains?
- How does MCTS combine parallel exploration with sequential reasoning depth?
- When is GPT model interpretation most likely to diverge from user intent?
- How can hidden test partitions detect constant predictions that generalize?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- Does policy entropy collapse in formal reasoning produce the same outcome in social reasoning?
- Does stable entropy in policy training actually guarantee stable reasoning behavior?
- What explains the gap between perplexity performance and actual reasoning capability?
- Why do current speech benchmarks fail to measure reasoning over audio?
- Why does outcome supervision fail for long reasoning chains?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- How does reinforcement learning differ from chain-of-thought distillation?
- How does RL compress reasoning path diversity during training?
- Can base models spontaneously produce reasoning traces without any RL training?
- How do extrapolative and contextual generalization measure RL reasoning gains?
- Can theory of mind models generalize across structurally similar scenarios?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- Can we predict out-of-distribution generalization without access to downstream tasks?
- What happens when models optimize specifically against CoT monitors?
- How much of Occamy's result comes from training versus the base model?
- How does demonstration coverage in context examples determine operation generalization?
- Does next-token prediction actually explain how human thought works?
- Does the token prediction framing actually capture what human reasoning does?
- Can standard next-token prediction capture complex multi-step human reasoning directly?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
DataAlchemy provides the empirical confirmation: imitation fails under distribution shift because no schema matches
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
distribution-bounded CoT is neither sufficient (fails under shift) nor necessary (in-distribution performance may not require the chain)
-
Can models pass tests while missing the actual grammar?
Do language models succeed on grammatical benchmarks by learning surface patterns rather than structural rules? This matters because correct outputs may hide reliance on shallow heuristics that fail on novel structures.
same pattern: surface patterns work in-distribution, fail under structural change
-
Does training data format shape reasoning strategy more than domain?
What explains why models trained on multiple-choice data reason differently than those trained on free-form text? The research isolates format and domain effects to measure which one matters more.
format-dependency is part of distribution-boundedness: changing the format is a distribution shift
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Hierarchical Reasoning Model
- Break the Chain: Large Language Models Can be Shortcut Reasoners
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Chain of Thoughtlessness? An Analysis of CoT in Planning
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
Original note title
cot reasoning is distribution-bounded — effectiveness degrades predictably with distributional discrepancy