Does longer reasoning actually mean harder problems?
Do chain-of-thought trace lengths reliably reflect problem difficulty, or do they primarily indicate proximity to training examples? Understanding this matters for designing effective scaling heuristics.
A prevailing assumption: longer reasoning traces indicate more thinking effort, therefore more complex problems should produce longer traces. Controlled experiments undercut this completely.
Training transformer models from scratch on derivational traces of the A* search algorithm — where problem complexity is precisely controllable and verifiable — reveals the decoupling:
- On in-distribution problems, trace length shows some alignment with difficulty
- On trivially simple problems (free-space mazes without obstacles), models often produce excessively long traces and sometimes fail to produce solutions
- On out-of-distribution problems, trace length and complexity become entirely decoupled — no correlation
The interpretation: intermediate token sequence length reflects approximate recall from the training distribution, not problem-adaptive computation. When a problem is close to training examples, the model retrieves a matching schema whose length reflects the training data's length distribution for that problem type. When a problem is far from training, the model has no calibrated schema to retrieve — trace length becomes arbitrary.
This challenges the entire anthropomorphic framing of "thinking time." When DeepSeek-R1 or similar models produce long chains, the conventional interpretation is that the problem is hard and the model is "working through it." The A* evidence suggests the length may primarily indicate how close the problem is to training distribution, not how much genuine computation is occurring.
The practical implication: trace length is not a reliable proxy for problem difficulty. Length-based scaling heuristics (add more tokens for harder problems) may be calibrating to the wrong signal. Does more thinking time always improve reasoning accuracy? supports this: more tokens do not reliably help after a certain point.
This also deepens Does chain-of-thought reasoning reveal genuine inference or pattern matching?: if trace length reflects training distribution proximity, then even the amount of imitation is calibrated to training similarity, not actual inferential needs.
Inquiring lines that read this note 166
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What explains the gap between benchmark scores and true reasoning capability?- Can benchmarks designed for shortcut learning detect heuristic override failures?
- Does the Heuristic Override Benchmark measure enumeration or world knowledge?
- Should benchmark evaluations use multiple prompt formulations for difficult tasks?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Why do short interaction benchmarks fail to predict long horizon performance?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- Can memorization inflate apparent capability on benchmarks with available solutions?
- Should long-context evaluation measure the coupled system?
- Can explicit constraint statements override the dominance of surface heuristics?
- Does the heuristic dominance ratio vary predictably across model architectures?
- Why do models automatically adjust reasoning length to problem difficulty?
- Why do simple math problems get worse with longer reasoning chains?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Are difficult tasks more monitorable because reasoning externalization becomes necessary?
- Do reasoning failures stem from strategy or from calculation breakdown?
- How do reasoning-related features behave when trained on near-impossible problems?
- Why does target probability matter more than task logical complexity?
- Does longer reasoning always improve model accuracy on complex tasks?
- How does task simplification affect analogical reasoning patterns?
- Why do simple length heuristics outperform sophisticated semantic methods?
- Are correct reasoning traces measurably shorter than incorrect ones?
- Are reasoning traces really reasoning or just stylistic imitation of human thought?
- What linguistic markers distinguish longer incorrect traces from correct ones?
- What makes a reasoning trace causally sufficient versus merely stylistically plausible?
- Why do correct reasoning traces appear shorter than incorrect ones?
- Can concise reasoning traces match verbose explanation accuracy?
- Why do shorter correct reasoning traces contain fewer failed branches?
- Why are correct reasoning traces consistently shorter than incorrect ones?
- Can reasoning traces serve purposes beyond producing the final answer itself?
- What saliency patterns distinguish successful from failed chain-of-thought reasoning?
- Why do correct reasoning traces tend to be shorter than incorrect ones?
- Why do we measure reasoning quality by reading visible chains?
- Why do reasoning traces resemble mimicry rather than verified problem-solving?
- Should benchmarks measure trace length or whether constraints were actually satisfied?
- How does trace coherence differ from valid mathematical proof in practice?
- How does trace coherence differ from trace validity in reasoning?
- Do correct reasoning traces tend to be shorter than incorrect ones?
- What makes some sentences in reasoning traces have disproportionate causal influence?
- What metric distinguishes deep reasoning from superficial information propagation?
- Do shorter correct reasoning traces contain more thought anchors than longer ones?
- Why do correct reasoning traces stay shorter than incorrect ones?
- Do longer chain-of-thought traces improve interpretability or just performance?
- How much of a reasoning trace is actually redundant or unnecessary?
- What distinguishes genuine capability gains from coherent but invalid reasoning traces?
- Why do reasoning traces persuade users without improving their accuracy?
- What makes a thinking trace take information shortcuts?
- What makes reasoning traces effective or ineffective for solving problems?
- Why are shorter reasoning traces more reliable than longer correct ones?
- Is the structure of reasoning traces learned as a shared stylistic convention?
- How do reasoning traces serve as hypotheses about decision processes?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- What causes reasoning loops and distraction in models with very long traces?
- What makes diffusion chain-of-thought reasoning qualitatively different from sequential chain-of-thought?
- Why do top performers produce shorter chains of thought in their strongest domains?
- How often do papers treat chain-of-thought as interpretability incorrectly?
- Why does chain-of-thought fail when problems lack matching training schemata?
- What happens to chain-of-thought performance across distribution shifts?
- Why do more capable models prefer shorter chains of thought?
- How does chain-of-thought training change higher layer computations?
- What structural properties define effective long chain-of-thought reasoning?
- How do exemplar properties affect the brittleness of chain-of-thought prompting?
- Why does chain-of-thought prompting fail to fix length-induced reasoning degradation?
- How do longer reasoning chains create vulnerability to attacks?
- What three factors actually drive chain of thought performance improvements?
- How does chain-of-thought length affect attention to constraint tokens?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- When is detailed step-by-step reasoning actually counterproductive for solving a problem?
- How much does chain-of-thought reasoning narrow the decompression gap?
- How does backtracking capability address error compounding in chain-of-thought reasoning?
- Why does extended chain-of-thought reasoning fail to improve numerical optimization performance?
- How does interaction horizon differ from chain-of-thought depth?
- Are chain-of-thought traces anthropomorphizing how AI models really reason?
- Can chain-of-thought traces harm rather than help user understanding?
- Does chain-of-thought accuracy degrade with longer reasoning traces?
- What makes some bottlenecks invisible to chain-of-thought training?
- How brittle are chain-of-thought exemplars across order and complexity?
- How does critique fine-tuning on one problem unlock broader reasoning?
- How can one training example improve reasoning across thousands of unseen problems?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How does capability evaluation differ from alignment evaluation in difficulty?
- How does distributional distance from pre-training relate to model difficulty?
- Does partial trace guidance work better than curriculum learning for hard problems?
- How do transformers generate harder solutions when mostly trained on easier problems?
- What makes preventative lessons from failures more valuable than success patterns?
- How do difficulty metrics relate to the true value of training examples?
- Why does SFT fail when expert demonstrations are too long for small models?
- Does training on curated solutions transfer to unseen problem types?
- Can solution traces substitute for process-level reward signals in math reasoning?
- Why does outcome supervision fail for long reasoning chains?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- What makes certain bond distributions more learnable than others?
- Can mathematical reasoning improvements transfer across problem subdomains?
- How do gradient descent iterations at inference compare to chain-of-thought reasoning chains?
- How does difficulty level change whether extended thinking provides genuine reasoning signal?
- Why do longer reasoning chains signal hesitation rather than depth?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- Can memorization scores diagnose where reasoning chains become unreliable?
- Do linearized traces genuinely expand exploration beyond standard chain-of-thought?
- Why do longer reasoning chains explore like tourists instead of scientists?
- How do task difficulty and skill type interact in model performance?
- Why does exemplar performance vary across order complexity diversity and style?
- Can scaling data alone solve performance gaps on long-tail concepts?
- What determines the finite chain length where robustness improvements plateau?
- Why do models overthink easy problems and underthink difficult ones?
- Does task difficulty alone determine how many thinking tokens a model should use?
- Can conditioning generation on difficulty probes reduce overthinking on simple tasks?
- When does extended thinking hurt performance on easier problems?
- Why do richer mental representations sometimes fail to predict better outcomes?
- Which RAG sub-decisions are actually pattern matching versus reasoning intensive?
- How does instance novelty rather than chain length explain reasoning failure?
- How do causal chains enforce long-horizon length differently than instruction-based tasks?
- Why do unresolved items cluster in structured patterns rather than randomly?
- How should inference budget adapt based on problem difficulty?
- Can test-time scaling work through retrieval rather than reasoning?
- How can systems estimate problem difficulty to allocate compute dynamically?
- Can event boundaries be identified from statistical regularities without understanding events?
- What makes a causal abstraction more transferable than a generic heuristic?
- Why does mixing reasoning traces from different teachers destabilize learning?
- Can deliberate corruption of reasoning traces harm out of distribution generalization?
- Do corrupted reasoning traces teach something different than pure success traces?
- Why does failed step fraction predict reasoning quality better than trace length?
- Why are incorrect reasoning traces longer than correct ones?
- Does trace length actually reflect problem difficulty or training proximity?
- Why do shorter confident reasoning traces fail on out-of-distribution problems?
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- How does confidence filtering improve selection of reasoning traces?
- What makes some reasoning traces better supervision than others despite equal accuracy?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- Can problem structure and representation format be mismatched intentionally?
- How do smaller models respond to longer reflection prompts?
- How do input length and context size separately affect reasoning quality?
- When does sequential chain-of-thought dramatically beat parallel voting approaches?
- What makes parallel thinking more efficient than sequential chains?
- What makes a problem fundamentally sequential versus parallelizable?
- When are multiple independent attempts more valuable than depth?
- Why do single-chain length and parallel exploration affect reasoning differently?
- What makes recursive depth more effective than parametric depth for puzzles?
- How sensitive is analogical reasoning emergence to training data and scale?
- Why do evaluation habits hide safety-critical challenges from view?
- How do inherited evaluation habits obscure failures that matter most?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do correct reasoning traces contain fewer tokens?
In o1-like models, correct solutions are systematically shorter than incorrect ones for the same questions. This challenges assumptions that longer reasoning traces indicate better reasoning, and raises questions about what length actually signals.
the within-distribution case: correct traces are shorter because they found the right schema quickly; this note explains the mechanism
-
Does more thinking time always improve reasoning accuracy?
Explores whether extending a model's thinking tokens linearly improves performance, or if there's a point beyond which additional reasoning becomes counterproductive.
practical consequence: tokens past the threshold reflect distribution mismatch, not useful computation
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
trace length is another dimension of imitation: how much training data looks like this problem
-
Does extended thinking actually improve reasoning or just increase variance?
When models think longer, do they reason better, or do they simply sample from a wider distribution of outputs that happens to cover correct answers more often? This matters because it determines whether test-time compute is genuinely scaling reasoning capability.
complementary: extended thinking broadens output distribution, not reasoning quality; trace length is part of this variance
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models
Original note title
cot trace length reflects training distribution proximity, not problem difficulty