SYNTHESIS NOTE
Topics›Looped Models›this note

Can reasoning systems scale faster by exploring parallel paths instead?

Current recursive reasoning models refine a single latent trajectory deeply, which is slow. Could sampling multiple trajectories in parallel achieve better reasoning with lower latency, and would that scale differently than serial refinement?

Synthesis note · 2026-05-28 · sourced from Looped Models

Recursive Reasoning Models (RRMs) increase reasoning capability by iterating a shared transition function over a latent state — more iterations means more "thinking" without extending the output sequence. This is depth scaling, and it decouples reasoning depth from both parameter count and output length. GRAM (Generative Recursive reAsoning Models) argues this is only half the story: depth alone is insufficient because a single refinement path can become trapped in a suboptimal trajectory, and many problems have ambiguity or multiple valid solutions that a single converging path cannot represent.

The structural claim is that future recursive reasoners should be not only deep (repeated refinement) but also wide (maintaining and exploring multiple latent trajectories in parallel). GRAM operationalizes width by turning the latent transition stochastic and sampling several trajectories simultaneously. Crucially, width sidesteps the latency penalty that depth-only scaling incurs: sampling N trajectories runs in parallel, whereas adding N refinement steps is serial and accumulates wall-clock time.

This reframes the inference-scaling design space for latent architectures. It mirrors at the latent-state level what parallel-vs-sequential debates established at the token level — since Why does parallel reasoning outperform single chain thinking?, breadth often beats depth under a fixed budget because independent paths sample the solution distribution rather than inflating variance along one path. GRAM brings that lesson inside the recurrent block, where prior work like Can models reason without generating visible thinking tokens? had only scaled depth. The counterpoint to watch: since Can parallel architectures solve inherently sequential problems?, width cannot substitute for depth on inherently serial problems — the two axes are complements, not interchangeable knobs. Why it matters: it gives latent reasoning a second, latency-cheap scaling dimension and explains why deterministic RRMs underperform on multi-solution tasks.

Inquiring lines that read this note 141

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Can intelligent routing over smaller models outperform scaling a single large model? How should items be represented and indexed in recommenders? How does decomposing tasks improve reasoning and prevent failure propagation? Can models improve accuracy without degrading reasoning quality? Can reasoning scale in latent space without tokens? Does RL create genuinely new reasoning capabilities or refine existing ones? How should test-time compute scaling work in agentic systems? Can inference-time compute effectively substitute for model scale? Can parallel reasoning outperform sequential reasoning under fixed token budgets? Does model confidence reliably signal actual accuracy in practice? What reasoning architectures enable models to solve complex problems efficiently? How should retrieval systems handle complex multi-step reasoning? How do surface patterns enable correct outputs but reduce robustness? How can evolutionary algorithms maintain diversity during solution search? How does misalignment propagate through agent communication networks? How should inference compute be allocated based on problem difficulty? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How does policy entropy collapse constrain scaling of reasoning-focused RL? When do multi-agent systems outperform single frontier models? How does self-revision in reasoning models affect accuracy and confidence? Can multi-agent systems avoid converging on false agreement without deliberation? How do capability benchmark scores systematically misrepresent true model abilities? Is reasoning capability latent in base models or created by post-training? What is the relationship between thinking tokens and reasoning accuracy? What causes reasoning models to fail or wander off track? How much does training format versus domain influence reasoning? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do soft reasoning mechanisms explore multiple paths without explicit training? What makes step-level supervision effective for complex reasoning traces? What role does sparsity play in model behavior and scaling decisions? What types of diversity prevent reasoning systems from collapsing? How does reasoning length affect model performance across different tasks? How can reward models capture diverse human preferences without excluding minority populations? How do neural networks achieve compositional generalization at scale? How much do training data properties shape model reasoning?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 113 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

reasoning systems should scale in width by sampling parallel latent trajectories not only in depth