Nobody taught AI models how far to reason, yet their depth and length limits appear anyway, as side effects of design.
How do depth and length constraints emerge in latent reasoning without explicit training?
This explores how models end up with limits on how deep they reason (how many internal steps) and how long they reason (how many tokens) when nobody trained those limits in directly, and what the corpus says about where those limits come from.
This explores where a model's limits on reasoning depth and reasoning length come from when no one trained those limits in directly. The corpus has no single paper on that exact question. Read together, though, several notes suggest an answer: these limits are mostly side effects. Some come from the architecture, some come from the reward signal, and some were already sitting in the model's internal representations.
Start with depth. Looped models run the same layers over and over instead of stacking more of them, and they can know when to stop without being taught to. The internal state settles down, and that convergence works as a built-in stopping signal Can models learn by looping instead of growing larger?. In other words, the model decides how deep to go based on when its own computation stops changing. Ouro builds this looping into pretraining itself, and its small models match much larger ones. Its intermediate loop states also line up closely with its final answers, so the depth it uses tracks real work rather than filler Can reasoning be learned during pretraining rather than after?. Two other results show this kind of iterated hidden computation can compete with spelled-out reasoning. One is a looped Transformer that matches explicit chain-of-thought on grade-school math Can latent reasoning close the scaling gap with explicit chain-of-thought?. The other is a tiny 150M-parameter model that takes on ARC-AGI puzzles at a fraction of a cent per task Can latent reasoning match chain-of-thought cost efficiency without verbalizing?.
Length shows the clearest pattern. Accuracy follows an inverted U: it rises as reasoning chains get longer, peaks, then falls. Harder tasks push the best length up, and more capable models push it down. RL training drifts toward shorter chains on its own as models improve, even though nobody rewards brevity Why does chain of thought accuracy eventually decline with length?. The surprising part is how simply the model stores this. Verbose and concise reasoning sit in different regions of the model's activations, and one direction separates them. A single vector taken from 50 example pairs cuts reasoning length by two-thirds without retraining or losing accuracy Can we steer reasoning toward brevity without retraining?. That fits a broader finding: base models already contain reasoning ability that post-training selects rather than creates Do base models already contain hidden reasoning ability?. Length preferences may be one more thing that was already there, waiting to be drawn out.
Not every length limit helps the model, though. Reasoning accuracy falls from 92% to 68% with just 3,000 tokens of padding, far below the context window's capacity, and chain-of-thought doesn't fix it Does reasoning ability actually degrade with longer inputs?. That ceiling wasn't designed by anyone, and it isn't the model being efficient. It's a weakness that emerged during training.
Here is the part you might not have expected: emergent depth and emergent length may be the same idea seen from two angles. A looped model stops when its state converges. An RL-trained model shortens its chain when the extra steps stop earning reward. In both cases the model reasons until more steps stop paying off, and nobody had to write that rule down.
Sources 8 notes
Models that re-apply layers in recurrent depth outperform larger feedforward networks on reasoning tasks. This works because recursion enables state tracking and compositional generalization that parameter scaling alone cannot achieve, with convergence signals providing natural halting.
Ouro's 1.4B–2.6B models match 12B baselines by performing reasoning during pretraining via iterative latent loops, not by storing more knowledge. Their intermediate latent states align strongly with final outputs, making them more faithful than divergent chain-of-thought traces.
LOTUS, a looped Transformer trained with gold CoT supervision at each latent position, closes the latent-to-explicit gap on GSM8K and cuts thought-phase latency by 2.5× to 6.9×. Latent reasoning need not be verbalized token-by-token to remain legible and competitive.
A 150M-parameter model combining in-context demonstrations with iterative latent computation reached 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, surpassing previously reported cost-accuracy tradeoffs. The approach separates learning (via demonstrations updating recurrent memory) from reasoning (via iteration in hidden space) without generating intermediate tokens.
Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.
Show all 8 sources
Activation-Steered Compression extracts a single vector from 50 paired examples to reduce chain-of-thought length by 67% while maintaining accuracy and achieving 2.73x speedup. The method is training-free and generalizes across model sizes and domains.
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
FLenQA shows reasoning accuracy drops from 92% to 68% at just 3000 tokens of padding, far below context window capacity. The degradation is task-agnostic, uncorrelated with language modeling performance, and persists even with chain-of-thought prompting.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- Hierarchical Reasoning Model
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models