Does RL training follow a predictable two-phase learning sequence?
This explores whether reinforcement learning exhibits consistent phases where basic execution skills must consolidate before strategic reasoning emerges. Understanding this sequence could reveal bottlenecks in scaling reasoning capabilities.
Across eight text-only and vision-language models, RL training reveals a consistently two-phase dynamic. In the first phase, the learning bottleneck is procedural correctness — a single calculation error invalidates an entire solution, creating powerful gradient signal that compels mastery of low-level execution tokens (arithmetic, variable substitution, formula application). In the second phase, the bottleneck shifts to strategic planning — exploring and mastering high-level planning tokens (deduction like "we can use the fact that," branching like "let's try a different approach," backtracing like "but the problem mentions that").
The phases are not mutually exclusive. Procedural refinement continues throughout training. But the primary driver of marginal performance gains shifts to strategic planning. This is why the "aha moment" phenomenon appears when it does — it represents the discovery and internalization of high-level reasoning strategies, which only becomes the active learning frontier after procedural skills are consolidated.
The entropy dynamics tell the same story. Planning tokens show increasing strategic diversification over training — the model explores new ways to combine established skills. Execution tokens show stable conditional entropy — once arithmetic is mastered, there's little incentive to find diverse ways to perform it. The performance improvement comes from discovering new combinations of established skills, which is the core function of planning.
This insight exposes a core inefficiency in algorithms like GRPO that apply optimization pressure uniformly across all tokens. If the learning frontier is in planning tokens but gradient signal is diluted across execution tokens, optimization is wasteful. HICRA addresses this by concentrating optimization on planning tokens, achieving significant performance gains.
The connection to existing insights is illuminating. Since Which sentences actually steer a reasoning trace?, HICRA's planning tokens are likely the same phenomenon identified from a mechanistic perspective. The two-phase dynamic also explains why Do reasoning cycles in hidden states reveal aha moments? — the graph structure reflects the transition from procedural execution (local structure) to strategic planning (global topology).
Inquiring lines that read this note 140
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance help or harm professional skill development?- Which AI interaction patterns preserve learning while which ones degrade skill formation?
- Does constraining AI access during early task phases preserve skill formation?
- Can explicit goal state scaffolding at inference time transfer to autonomous tracking through training?
- Why does early experience provide better warm-starts for downstream reinforcement learning?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
- How does sliding the start state backward create informative learning signals?
- What limits RL's ability to scale for reasoning at training time?
- Which recipe choices determine the asymptotic ceiling in RL training?
- How does behavior cloning reduce complexity before RL training in rerankers?
- What distinguishes RL that creates new capabilities from RL that merely teaches timing?
- How does trajectory burstiness compare to other structural properties that shape emergent capabilities?
- What structural differences emerge between early generic skills and later meta-strategy skills?
- How does post-training shift models from passive prediction to on-policy action?
- Does RL training activate latent meta-learning capacity or create it from scratch?
- What training duration is actually needed for RL to expand capabilities?
- What capacity threshold determines whether RL teaches activation versus shortcut learning?
- How does pretraining determine what RL can later teach a model?
- How does model scale affect anticipatory behavior in structured training?
- Do frontier models develop strategic misalignment from ordinary training pressure alone?
- How does in-context learning trigger phase transitions in model behavior?
- How do training objectives shape what a world model actually learns?
- How do developmental curriculums emerge from learning progress signals?
- How should guidance levels adapt as the model's capability boundary shifts?
- Why does imitation learning alone plateau without outcome-based refinement?
- How do complete multi-turn trajectories differ from isolated task examples?
- What training interventions could close the perception-action gap?
- Does the productive difficulty band ever stabilize during training?
- How does the optimal difficulty band shift as the model's capabilities improve during training?
- What makes exploration a verifiable and measurable training objective?
- What distinguishes surface mechanisms from the training regimes that produce them?
- How does scaffolding unstable mechanics improve reinforcement learning for search?
- How do learning dynamics on one example shift predictions on other responses?
- Does therapy environment difficulty calibration affect RL policy learning quality?
- Can outcome-based rewards fully replace per-step likelihood in diffusion RL training?
- How do evaluative versus directive signals differ in next-state training?
- Why do next-turn reward objectives fail to encourage multi-turn goal progress?
- Can continuous spectrum training outperform sequential SFT-then-RL stages?
- What scaling properties emerge from RL training dynamics beyond verification?
- How does curriculum learning prevent instability in social-emotional RL training?
- Why do single-turn RL methods fail to generalize to multi-turn tasks?
- How should multi-objective post-training balance competing behavioral goals?
- How does temporal anchoring maintain learning signals when preference gaps collapse?
- Why does alternating RL training stabilize learning better than simultaneous updates?
- What behavioral changes occur during reward learning training?
- Does task ordering affect multi-task reinforcement learning outcomes?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- Can in-context learning replicate the timing effects that RL teaches models?
- Can meta-reinforcement learning explain why this bias pattern emerges rationally?
- Can RL teach when to use reasoning versus when to respond directly?
- Does reinforcement learning learn optimal per-turn reasoning discipline?
- How do residual connections and layer norm stabilize training in deep RL?
- How does reinforcement learning differ from chain-of-thought distillation?
- Does RL refine existing knowledge or discover entirely new capabilities?
- How does RL compress reasoning path diversity during training?
- Why do models follow a two-phase pattern of procedural then strategic learning?
- Can models learn both what and how to study through reinforcement learning?
- Does format-based pretraining determine how models respond to reinforcement learning?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Can out-of-distribution tests expose memorization in reinforcement learning fine-tuned models?
- Why does prolonged RL discover strategies absent from any base model sample?
- Can reinforcement learning fix the reasoning gaps that supervised fine-tuning misses?
- Why does RL behavior differ between standard reasoning tasks and complex planning domains?
- Does reinforcement learning teach models how to reason or when to reason?
- Does RL primarily teach when to use reasoning or how to reason?
- What does RL post-training actually teach reasoning systems?
- What makes supervised fine-tuning worsen RL exploration later?
- Does RL training redirect self-doubt into productive gap analysis?
- When does reinforcement learning actually produce true reasoning gains in models?
- How does imitation pretraining followed by RL exploration compare to either method alone?
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- What distinguishes high-signal prompts from low-signal ones in RL training?
- Does RL teach models new reasoning or just better timing?
- Does reinforcement learning require sufficient pretraining to be effective?
- Should larger compute budgets allocate more resources to reinforcement learning?
- Does the reinforcement learning improvement rate depend on model initialization?
- How does entropy collapse in reinforcement learning differ from entropy maintenance in graph reasoning?
- Does policy entropy collapse represent the main bottleneck in reasoning-focused RL scaling?
- Does policy entropy collapse limit how many iterations of reasoning training work?
- How does policy entropy during training affect search discipline during inference?
- Why does policy entropy collapse predict sigmoid saturation points?
- What happens to model reasoning when policy entropy collapses during RL?
- Why do high entropy tokens carry most of the learning signal in RL?
- How does representational convergence differ from policy entropy collapse in iterative training?
- How do high-entropy tokens concentrate reinforcement learning's effect?
- How does on-policy entropy recognition differ from training-time entropy collapse?
- Why does policy entropy collapse when scaling RL for reasoning?
- Can entropy regularization or critique models prevent search strategy collapse during RL training?
- What causes policy entropy collapse in reasoning-focused reinforcement learning?
- How does policy entropy collapse constrain zero RL scaling for reasoning?
- Does stable entropy in policy training actually guarantee stable reasoning behavior?
- Does post-training collapse policy entropy more than base model sampling?
- Why does prolonged RL with entropy control beat base models at all pass@k levels?
- Why must procedural skills consolidate before strategic reasoning can develop?
- What makes bilevel metacognition architectural rather than emergent in current systems?
- Do depth thresholds correspond to transitions between procedural and strategic learning?
- Do thought anchors correspond mechanistically to planning tokens in RL?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- What theoretical argument connects iterative DPO dynamics to online RL learning?
- What specific properties of online RL does iterative DPO actually preserve?
- Does iterative DPO preserve the same misalignment mechanisms as online reinforcement learning?
- Does iterative online DPO fidelity match true reinforcement learning for safety research?
- Do outcome-only reward signals miss step-level errors that compound later?
- What makes Effective Rank Acceleration a stable training signal for dual-channel incentives?
- How do reward signals in RLVR interact with pretraining biases?
- What makes advantage shaping more stable than reward shaping for tool training?
- Why does outcome-only reinforcement learning need more than double the tokens to train agents?
- How does dual-rate learning separate episodic and procedural memory in neural networks?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- Does grokking in modular arithmetic follow the same three-phase learning trajectory?
- How do complementary learning systems explain the need for fast and slow consolidation?
- Can architectural changes like decoupling intent understanding help overcome next-turn reward limitations?
- How does next-turn reward optimization contribute to agent passivity?
- Can early experience replace external rewards as a learning signal?
- Can RL training teach models when to activate reasoning versus when to skip it?
- Does RL training actually restore the critical thinking that reasoning models lose?
- How does policy initialization with sub-policies enable emergent thinking?
- Does targeting the edge of competence during RL pretraining unlock true reasoning gains?
- How do two-phase training dynamics explain reasoning emergence?
- How does memory folding enable agents to reconsider strategies mid-task?
- Does operator-conditioned memory let search compose learned behaviors more effectively?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Which sentences actually steer a reasoning trace?
Can we identify which sentences in a reasoning trace have outsized influence on the final answer? Three independent methods converge on a surprising answer about planning and backtracking.
converges: planning tokens in HICRA likely correspond to thought anchors
-
Do reasoning cycles in hidden states reveal aha moments?
What if the internal loops in model reasoning—visible in hidden-state topology—correspond to the reconsidering moments that happen during reasoning? This note explores whether graph cyclicity captures a mechanistic signature of insight.
extends: the two-phase dynamic explains how graph topology evolves during training
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
reframes: entropy collapse may be acceptable for execution tokens but catastrophic for planning tokens
-
Does RL teach reasoning or just when to use it?
Does reinforcement learning in thinking models actually create new reasoning abilities, or does it simply teach existing capabilities when to activate? This matters for understanding where reasoning truly emerges.
deepens: the "when" is specifically about planning tokens; execution tokens are "how"
-
What happens inside models when they suddenly generalize?
Grokking appears as an abrupt shift from memorization to generalization. But is the underlying process truly discontinuous, or does mechanistic analysis reveal continuous phases we can measure and predict?
analogous phased development: grokking's memorization-then-circuit-formation parallels the procedural-then-strategic progression; both show that generalization requires passing through a consolidation phase before higher-order structure emerges
-
Can language modeling close the knowing-doing gap in AI?
Current LLMs reason well but act poorly in interactive tasks, while RL agents act well but cannot explain themselves. Can reformulating decision-making as language modeling with environmental feedback bridge this fundamental split?
TiG operates on the same procedural-vs-strategic axis HICRA identifies, but at the architectural level: language-as-policy refined by RL preserves declarative reasoning while building procedural competence — HICRA's two-phase dynamic predicts the order TiG observes during training
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Emergent Hierarchical Reasoning In LLMs Through Reinforcement Learning
- Reinforcement Learning with Rubric Anchors
- Sharpening Tax in Post-Training
- RAGEN-2: Reasoning Collapse in Agentic RL
- The Art of Scaling Reinforcement Learning Compute for LLMs
- From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- Understanding Reasoning from Pretraining to Post-Training
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Original note title
rl training exhibits a two-phase dynamic where procedural consolidation precedes strategic planning exploration