Thinking longer doesn't always help an AI: past a certain point, extra reasoning can wreck answers it would otherwise get right.
When does richer information actually harm decision quality?
This explores when giving an AI more to work with (more reasoning steps, longer chains of thought, harder training material) makes its decisions worse instead of better. The corpus answers this mostly for a model's own self-generated reasoning, not for human decision-makers drowning in data.
This explores when more information makes AI decisions worse instead of better. In this collection, the 'more' is mostly the model's own thinking: longer reasoning traces, extra tokens and added deliberation. The pattern is consistent. More helps up to a point and then hurts. In one study, raising a model's thinking budget from about 1,100 to 16,000 tokens dropped benchmark accuracy from 87.3% to 70.3% Does more thinking time always improve reasoning accuracy?. The same rise-then-fall curve shows up for chain-of-thought length in general Why does chain of thought accuracy eventually decline with length?, and the overview note on overthinking gathers the evidence in one place When does thinking too much actually hurt reasoning?.
The mechanism is the surprising part. Extra reasoning doesn't just stop helping. It actively destroys good answers. Researchers who tracked 'flip events' found that models often reach the correct answer, keep thinking, talk themselves out of it, and write a wrong one over the top Why does more reasoning sometimes make models worse?. Here, richer deliberation works like a bad committee meeting: more chances to second-guess and more output variance, with no new evidence coming in. A related finding makes this sharper. Chain-of-thought examples that are logically *invalid* help almost as much as valid ones Does logical validity actually drive chain-of-thought gains?. So a lot of what looks like 'more reasoning' is the shape of reasoning, not the substance. If the extra tokens carry little real content, they can't be relied on to fix mistakes, but they can still introduce new ones.
Whether extra thinking hurts depends on the reasoner, not on how much thinking there is. Untrained models use extended thinking to spiral into self-doubt. After reinforcement learning, the same mechanism becomes useful gap-checking Does extended thinking help or hurt model reasoning?. More capable models also do best with *shorter* chains, while harder tasks call for longer ones Why does chain of thought accuracy eventually decline with length?. So there's no single 'right amount'. That has led to tools that watch the reasoning as it happens. One uses the model's confidence patterns to tell overthinking from underthinking and steer between them Can confidence patterns reveal overthinking versus underthinking?. Another measures how many tokens reflect genuine internal revision versus filler Can we measure how deeply a model actually reasons?.
The counterpoint matters. Richer *external* feedback usually helps. A single score of 'good' or 'bad' throws away the *how to fix it* information that natural feedback carries Can scalar rewards capture all the information in agent feedback?. Models stuck on a performance plateau with numerical rewards can break through when given written critiques instead Can natural language feedback overcome numerical reward plateaus?. The real line is not rich versus sparse but signal versus noise. Training shows the same thing. Problems that are far too hard give the model mostly accidental successes, the model learns shortcuts from those flukes, and the shortcuts damage skills it already had Do overly hard RLVR samples actually harm model capabilities?.
The takeaway: in this corpus, more information hurts when it is self-generated or unlearnable, meaning more of the system's own deliberation, or more examples it can't actually learn from. It helps when it brings in something new from outside, like a critique that explains what went wrong. The collection doesn't yet have much on the human version of this question, such as decision fatigue or information overload for people using AI. That's a real gap, not something this answer can fill.
Sources 11 notes
Increasing thinking tokens from ~1,100 to ~16K reduced benchmark accuracy from 87.3% to 70.3%, revealing a non-monotonic relationship where models overthink easy problems and underthink hard ones.
Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.
Empirical studies demonstrate non-monotonic scaling in test-time reasoning: accuracy peaks at a critical thinking-token count, then declines sharply (87.3% to 70.3% as tokens scale from 1,100 to 16,000). Extended thinking inflates output variance and introduces self-revision errors rather than improving solution quality.
Tracking flip events shows that extra reasoning tokens don't just hit diminishing returns—they actively cause models to second-guess and overwrite previously-correct answers, making accuracy non-monotonic in trace length.
Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.
Show all 11 sources
Vanilla models use thinking mode counterproductively, inducing self-doubt that degrades performance. RL training reverses this, transforming the same mechanism into beneficial gap analysis. Training mediates reasoning quality, not just quantity.
ReBalance uses confidence variance and overconfidence as diagnostic signals to apply training-free steering vectors that reduce overthinking redundancy while promoting exploration during underthinking, improving accuracy across models from 0.5B to 32B parameters.
Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.
Natural feedback carries two orthogonal types of information: evaluative (how well an action performed) and directive (how it should change). Scalar rewards capture evaluation but discard directional specifics that token-level distillation can recover, making the two complementary rather than redundant.
Critique-GRPO shows that models stuck on performance plateaus can generate correct solutions when given chain-of-thought critiques, revealing that numerical rewards lack critical information about why failures occur and how to improve.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
- Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- Understanding and Mitigating Premature Confidence for Better LLM Reasoning
- When More is Less: Understanding Chain-of-Thought Length in LLMs