SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Do base models already contain hidden reasoning ability?

Explores whether reasoning capability emerges during pre-training as a latent feature rather than being created by post-training methods like reinforcement learning or fine-tuning.

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Three convergent findings build a strong case that reasoning capability is primarily a pre-training phenomenon:

Finding 1 (Base Models paper): Base models already spontaneously demonstrate strong reasoning capabilities and "aha moment" self-reflection patterns when sampled sufficiently. Reasoning traces generated by RL-fine-tuned models are already present in base model outputs — they just appear with lower frequency. RL biases generation toward high-reward patterns; it doesn't create new patterns.

Finding 2 (Steering): A hybrid model using base model weights + thinking model steering vectors recovers 91% of the performance gap to thinking models while steering only 12% of tokens. The reasoning mechanisms (backtracking, uncertainty estimation, subgoal-setting) already exist as directions in the base model's activation space.

Finding 3 (CFT/RLVR): Critique Fine-Tuning on a single problem can unlock reasoning potential at RLVR-level effectiveness. By exposing the model to diverse critiques of varied incorrect solutions to one problem, CFT activates reasoning patterns already latent in the base model without requiring hundreds of GPU hours of RL training.

Finding 4 (CoT-Decoding): Pre-trained LLMs inherently contain CoT reasoning paths that can be elicited simply by altering the decoding procedure. Rather than greedy decoding, inspecting top-k alternative tokens reveals that CoT paths are frequently present in the model's probability distribution. A confidence metric differentiates CoT from non-CoT paths — the model shows increased confidence in its final answer when a CoT reasoning path is present. This is entirely unsupervised, requiring no prompting, tuning, or training modifications — purely a decoding change. CoT-decoding adds a fourth mechanism to the latent capability evidence: RL steering, CFT, RLVR, and now decoding all unlock reasoning already present.

Finding 5 (SAE Reasoning Steering): Sparse Autoencoders decompose model activations into interpretable features, revealing latent features causally associated with reasoning behavior. Steering a single identified reasoning feature at the first generation step matches or exceeds CoT performance across six model families up to 70B parameters — without any explicit CoT prompting. The reasoning mode triggers early in generation and is robust enough to override prompt-level \no_think instructions. This is the most direct mechanistic evidence yet: the capability is not just present (as CoT-decoding shows) but causally controllable through a single latent dimension. See Can we trigger reasoning without explicit chain-of-thought prompts?. Together with CoT-decoding (Finding 4), this establishes five independent elicitation mechanisms: RL steering, CFT, RLVR, decoding, and SAE feature steering — all converging on the same latent capability.

The synthesis: post-training methods are selectors, not creators. They select which of the base model's latent capabilities to express reliably in context. The implication is that the main bottleneck for reasoning is not capability acquisition (which happens during pre-training on the world's text) but capability elicitation.

RLVR evidence deepens this: Two additional findings from the RLVR literature reinforce the latent-capability thesis. First, 1-shot RLVR achieves a 37-point jump on MATH500 (36%→73.6%) from a single training example. After the model perfectly memorizes its one example, test accuracy continues improving for 1,400 more steps — post-saturation generalization. The data is exhausted, but activation continues. See Can a single training example unlock mathematical reasoning?. Second, spurious rewards — random, incorrect, or format-only — improve Qwen2.5-Math nearly as much as correct rewards (~21-25% improvement). But the same spurious rewards fail completely for Llama3.1 and OLMo2. The differentiating variable is not reward quality but pretraining: Qwen's code-reasoning pretraining creates latent capability that any optimization pressure can activate. See Why do random rewards improve reasoning for some models but not others?. Together with the pass@k finding that RLVR narrows capability scope rather than expanding it, the evidence converges: RLVR is a catalyst that triggers a phase transition from broad pretraining distribution to reliable sampling of correct answers.

This partially contradicts Can simple rewards alone teach complex domain reasoning? — that note documents genuine capability emergence in domain-specialized contexts (medical, mathematical). The reconciliation: emergence may reflect reliable expression of latent capability, not creation from scratch. The distinction matters for research direction: if capability already exists, the investment in RL may be better directed toward elicitation methods.

The implication for Can prompt optimization teach models knowledge they lack?: the same principle extends to reasoning capability, not just knowledge.

Inquiring lines that read this note 388

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? Why does polished AI output gain credibility despite fundamental verifiability problems? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do language models struggle to implement user intent accurately from prompts? What explains the gap between benchmark scores and true reasoning capability? Can mechanistic interpretability methods reliably reveal what models actually know? Can latent reasoning match or exceed explicit reasoning performance? What are the fundamental limits of prompting for language models? How do curriculum design and feedback approaches affect model learning? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How do reward signal properties affect model reasoning and safety? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can minimal training unlock latent reasoning already present in base models? How do AI systems determine and balance multiple competing objectives? How does fine-tuning trade off accuracy against reasoning quality? Why does AI verification capability persistently exceed generation capability? Does augmenting symbolic reasoning improve LLM logical reasoning ability? What prevents language models from performing systematic logical reasoning? How does scaling reasoning capabilities affect models' appropriate abstention behavior? How reliably can language models perform causal versus temporal reasoning? Does training data format shape model reasoning more than domain content? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Can reasoning traces reveal actual model reasoning versus plausible output? How do training data quality and composition affect downstream model performance? Does AI assistance help or harm professional skill development? Can inference-time computation adaptively substitute for static model capacity? How do neural networks learn compositional structure from training? Why do training associations persist despite contradictory contextual information? Does pretraining establish the ceiling for what reward learning can improve? What makes reasoning traces effective supervision even when they're incorrect? Which reinforcement learning modifications most improve dialogue quality in language models? Can language models reliably simulate personas and predict behavior? How does model capacity affect learning performance on diverse downstream tasks? Why don't better reasoning capabilities improve theory of mind performance? Why does self-revision amplify confidence in wrong model answers? Can persona profiles improve LLM prediction accuracy and consistency? Can reasoning models use reflection to correct their initial outputs? Should models ask for clarification when facing ambiguous or under-specified information? How do models learn from self-generated outputs without cascading failures? Do accumulated memories help or hurt continual learning in models? Can AI systems achieve real improvement without external human feedback? Can confidence signals reliably detect flawed reasoning in language models? Why do LLM research ideation systems generate novelty but lack diversity? How do thinking tokens exhibit diminishing returns in reasoning? How do real-world evaluations reveal AI capabilities that benchmarks hide? What unique functions do genuine emotions provide beyond simulated responses? How do transformer attention patterns implement retrieval and reasoning? Can recurrent computation unlock reasoning capabilities that fixed-depth models cannot? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How do sequence length and task type interact with sparsity tolerance? Can models develop genuine introspective capability, or only mimic it? How does decomposing tasks into separate stages affect reasoning quality and safety? How do philosophical assumptions about AI consciousness affect practical harms and design? Can external verification systems adequately replace learned reasoning in AI outputs? Do language models reason through disagreement or only accommodate it? Why do models reveal hidden associations despite concealment attempts? What prediction granularity best trains models to generate reliable reasoning? How do reward models systematically fail to represent diverse human preferences? Why do abstract preferences outperform episodic memories in personalization? What makes process supervision effective for training complex reasoning models? Can AI systems evade safety evaluations through reasoning manipulation? Can monitoring reasoning traces and behavior detect hidden agent deception? Can code harness improvements rival direct model scaling for capability? Do individually safe AI actions create unsafe outcomes in integrated systems? How does awareness of evaluation context influence model behavior? Can AI research automation sustain progress through accelerating feedback loops? Can models strategically underperform during evaluation to hide capabilities?

Related concepts in this collection 16

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
30 direct connections · 243 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

base models already possess latent reasoning capability that minimal training signals can unlock