SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Does RL post-training create reasoning or just deploy it?

Investigates whether reasoning capability emerges during RL fine-tuning or already exists in base models. Matters because it reshapes how we build and optimize reasoning systems.

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Post angle — Medium/LinkedIn

The dominant story: DeepSeek R1, GPT-o1, and their successors acquire reasoning capability through RL post-training. RL teaches models to think step-by-step, to backtrack, to verify — capabilities they didn't have before.

The emerging counter-evidence is striking. A hybrid model using a base model's weights with a thinking model's deployment decisions — zero weight updates — recovers 91% of the performance gap to thinking models by steering only 12% of tokens. Base models already spontaneously produce reasoning traces identical to thinking model traces when sampled sufficiently. Single-problem CFT achieves RLVR-level reasoning gains. Activation-space vectors encoding "backtracking" and "uncertainty estimation" already exist in base model hidden states before any RL.

The reframe: pre-training is when reasoning capability is acquired; RL post-training teaches when to deploy it.

This is not a trivial distinction. "When" training is cheaper, less data-hungry, and less fragile than "how" training. If capability already exists, elicitation methods (structured tool-calling, steering vectors, targeted fine-tuning on single problems) become much more attractive than full RL pipelines.

The hook for readers: "We've been crediting the locksmith for the key."

Connections: Does RL teach reasoning or just when to use it?, Do base models already contain hidden reasoning ability?, Can modular cognitive tools unlock reasoning without training?

Inquiring lines that read this note 177

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished AI output gain credibility despite fundamental verifiability problems? Does pretraining establish the ceiling for what reward learning can improve? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can minimal training unlock latent reasoning already present in base models? Can latent reasoning match or exceed explicit reasoning performance? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Can mechanistic interpretability methods reliably reveal what models actually know? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Does training data format shape model reasoning more than domain content? How do training data quality and composition affect downstream model performance? How do curriculum design and feedback approaches affect model learning? How does fine-tuning trade off accuracy against reasoning quality? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Why do planning and grounding require opposing optimization strategies? What prevents LLMs from applying their reasoning knowledge to improve outputs? What makes reasoning traces effective supervision even when they're incorrect? What are the fundamental limits of prompting for language models? What prevents language models from performing systematic logical reasoning? How does model capacity affect learning performance on diverse downstream tasks? Can AI agents improve their skills through accumulated experience and reuse? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can artificial systems establish authority in domains requiring expert judgment? Can AI systems achieve real improvement without external human feedback? What design features sustain romantic bonds with AI companion systems? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Does intelligent routing among smaller models outperform training larger models? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? What makes process supervision effective for training complex reasoning models? Do accumulated memories help or hurt continual learning in models? Can models develop genuine introspective capability, or only mimic it? Which reinforcement learning modifications most improve dialogue quality in language models? Can reasoning traces reveal actual model reasoning versus plausible output? Should agents compress episodic memory or retain raw interaction histories? How do neural networks learn compositional structure from training? Can smaller specialized models match frontier models on key metrics? Can external verification systems adequately replace learned reasoning in AI outputs? Why do models reveal hidden associations despite concealment attempts? How do agents learn to distinguish valuable feedback from noise? Do individually safe AI actions create unsafe outcomes in integrated systems? How does awareness of evaluation context influence model behavior? How do AI systems determine and balance multiple competing objectives? Can code harness improvements rival direct model scaling for capability? Does AI assistance help or harm professional skill development?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 192 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

thinking models learn when not how — the case that rl post-training is a deployment optimizer not a capability creator