Why doesn't more AI 'thinking' always mean better answers — sometimes it even makes things worse?
Why do richer mental representations sometimes fail to predict better outcomes?
This explores why giving a model more internal 'thinking' — longer reasoning chains, more thinking tokens, or more elaborate latent representations — often fails to improve its answers, and sometimes makes them worse.
This explores why more elaborate internal reasoning in AI models, such as longer chains of thought, bigger thinking budgets or richer hidden representations, doesn't reliably produce better answers. The short version from the corpus: richness only helps when it's aimed at the actual bottleneck, and much of what looks like rich reasoning is something else in disguise.
The most direct evidence is that reasoning length follows an inverted U. Accuracy rises with chain-of-thought length up to a point and then falls. The best length grows with task difficulty but shrinks as models get more capable Why does chain of thought accuracy eventually decline with length?. In one measurement, pushing thinking from about 1,100 to 16,000 tokens dropped benchmark accuracy from 87% to 70%. Models overthink easy problems and underthink hard ones Does more thinking time always improve reasoning accuracy?. Extra thinking also isn't neutral. In untrained models it can turn into self-doubt that talks the model out of correct answers, and only RL training redirects the same mechanism into useful analysis Does extended thinking help or hurt model reasoning?.
A more surprising answer is that the elaborate trace often isn't doing the work it appears to do. Reasoning chains with deliberately invalid logic perform nearly as well as valid ones, which suggests models benefit from the *shape* of reasoning more than its content Does logical validity actually drive chain-of-thought gains?. Trace length tracks how familiar a problem is from training, not how hard it is. On unfamiliar problems the link between length and difficulty disappears entirely Does longer reasoning actually mean harder problems?. One study split chain-of-thought performance into three parts: output probability, memorization and real but noisy step-by-step reasoning that piles up errors with each added step What three separate factors drive chain-of-thought performance?. So a longer chain can mean more recall plus more chances to go wrong. Related work argues that base models already contain much of their reasoning ability, so post-training mostly *selects* it rather than building it. A more elaborate trace can't supply ability that wasn't there to begin with Do base models already contain hidden reasoning ability?.
The third reason is a mismatch between the representation and the problem. On vision tasks, long verbal reasoning *hurts* fine-grained perception, because the real limit is where the model directs its visual attention, not how much it can say Does verbose chain-of-thought actually help multimodal perception tasks?. Reasoning done in hidden vectors instead of words fails in its own way. When training only checks the final answer, the learning signal fades across the hidden steps and the hidden space drifts away from meaning, so the representation becomes richer in dimensions but emptier in content Why does latent chain-of-thought fail so easily in training?.
The counterexamples show when richness does pay off. A hierarchical planning model jumped from 18% to 73% success by giving each level of its hierarchy its own more abstract representation, built to match what that level has to judge Does each hierarchy level need its own latent space?. Likewise, models stuck on a plateau under a single numerical reward improve when given written critiques, a richer signal that says *why* they failed Can natural language feedback overcome numerical reward plateaus?. The takeaway you might not have expected: a rich representation helps when it carries information the task actually needs and is trained against the right target. Volume alone tends to give memorization and accumulated noise more room, or simply more chances to second-guess.
Sources 11 notes
Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.
Increasing thinking tokens from ~1,100 to ~16K reduced benchmark accuracy from 87.3% to 70.3%, revealing a non-monotonic relationship where models overthink easy problems and underthink hard ones.
Vanilla models use thinking mode counterproductively, inducing self-doubt that degrades performance. RL training reverses this, transforming the same mechanism into beneficial gap analysis. Training mediates reasoning quality, not just quantity.
Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.
Controlled A* maze experiments show trace length correlates with difficulty only in-distribution but decouples entirely out-of-distribution. Trace length primarily reflects recall of training schemas, not adaptive computation.
Show all 11 sources
A shift cipher study decomposed CoT into three independent factors: output probability alone swings accuracy from 26% to 70%, memorization matches pre-training frequency patterns, and genuine reasoning exists but accumulates error with each step. This resolves the reason-or-memorize debate by showing LLMs do both simultaneously.
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
Long rationales and text-token RL help reasoning but hurt fine-grained perception tasks because the actual bottleneck is visual attention allocation, not verbalization. Standard CoT optimization trains the wrong policy target.
Outcome supervision alone causes gradient attenuation along latent steps and lets the latent space wander without semantic grounding. Robust latent reasoning requires both dense trajectory supervision and space supervision that preserves geometric structure rather than compressing it.
H-JEPA improves Visual AntMaze planning from 18% to 73% success by giving each hierarchy level a distinct, more abstract latent space rather than sharing one. This per-level abstraction lets higher levels score candidate futures in abstract space better matched to goal-like objectives.
Critique-GRPO shows that models stuck on performance plateaus can generate correct solutions when given chain-of-thought critiques, revealing that numerical rewards lack critical information about why failures occur and how to improve.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- Base Models Know How to Reason, Thinking Models Learn When