SYNTHESIS NOTE
Topics›Conversation Agents›this note

Does extended thinking help or hurt model reasoning?

Explores whether activating thinking mode improves reasoning performance, and what role training plays in determining whether extended internal reasoning chains are productive or counterproductive.

Synthesis note · 2026-02-22 · sourced from Conversation Agents

The proactive critical thinking experiments reveal a striking interaction between training and inference-time reasoning. For vanilla (off-the-shelf) models, activating "thinking mode" — the extended internal reasoning chains used by models like Qwen3 — actually degrades performance on proactive critical thinking tasks. The extended thinking "appears to induce counterproductive self-doubt rather than useful analysis, leading to a clear drop in performance."

But after RL training on proactive critical thinking tasks, the same thinking mode becomes beneficial. Training fundamentally changes how models use their internal reasoning. This is not merely about more or less thinking — it is about the quality direction of thinking.

The finding connects to several established insights but adds a distinct mechanism:

Since Does RL teach reasoning or just when to use it?, RL manages the timing of reasoning. The proactive thinking result extends this: RL also manages the mode of reasoning — redirecting extended thinking from unproductive self-doubt toward productive gap analysis.

The SFT finding adds nuance: when SFT data is self-generated by the model, it "does not inherently enhance its capabilities" and may reduce output entropy, constraining the subsequent RL phase. This echoes Does policy entropy collapse limit reasoning performance in RL? — SFT-then-RL may face the same entropy collapse that pure RL faces, but through a different mechanism (entropy reduction from self-generated imitation rather than RL convergence).

The practical implication: extended thinking is not a universal good. It is a resource that can be directed productively or destructively, and the direction depends on training. "More thinking" applied to a model without the right training signal may systematically make things worse.

Inquiring lines that read this note 141

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do language models fail at sustained therapeutic relationships despite understanding techniques? How do reward signal properties affect model reasoning and safety? Can minimal training unlock latent reasoning already present in base models? Can latent reasoning match or exceed explicit reasoning performance? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Can reasoning models use reflection to correct their initial outputs? How does scaling reasoning capabilities affect models' appropriate abstention behavior? How do thinking tokens exhibit diminishing returns in reasoning? What prevents LLMs from applying their reasoning knowledge to improve outputs? What limits recursive self-improvement in autonomous AI systems? When does parallel reasoning outperform sequential reasoning with the same token budget? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? What makes reasoning traces effective supervision even when they're incorrect? How do interpretive frames override surface features in text comprehension? Can confidence signals reliably detect flawed reasoning in language models? What gaps exist between benchmark performance and real deployment outcomes? What prevents language models from performing systematic logical reasoning? How reliably can language models perform causal versus temporal reasoning? Why don't better reasoning capabilities improve theory of mind performance? Do persona-based approaches introduce systematic biases in user simulation? Can mechanistic interpretability methods reliably reveal what models actually know? Why does self-revision amplify confidence in wrong model answers? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How should retrieval strategies adapt to multi-step reasoning demands? What limits language model accuracy in evaluating ideas? Can reasoning traces reveal actual model reasoning versus plausible output? How do curriculum design and feedback approaches affect model learning? What unique functions do genuine emotions provide beyond simulated responses? Can LLMs distinguish between linguistic form and semantic meaning? What are the fundamental limits of prompting for language models? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? What makes process supervision effective for training complex reasoning models? How do users confuse explanation quality with actual system accuracy? Can AI systems achieve real improvement without external human feedback? How does awareness of evaluation context influence model behavior? How do AI systems determine and balance multiple competing objectives? Does AI assistance erode cognitive skills while inflating perceived competence?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 170 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl training transforms thinking mode from counterproductive self-doubt into beneficial proactive analysis — the same mechanism helps or hurts depending on training