Training changes an AI's weights slowly and permanently — but tricks run at use-time can boost it just as much, instantly and reversibly.
What distinguishes fast non-parametric loops from slow parametric weight updates?
This explores the difference between improving an AI system by changing its weights through training (slow, permanent, parametric) and improving it by adding loops, interventions, or scaffolding around fixed weights (fast, reversible, non-parametric), and what each approach can and can't do.
This explores two ways to make a model better: retrain its weights, or leave the weights alone and change what happens around them at run time, through repeated passes, interventions on its internal activity, or outside scaffolding. The corpus suggests the line between the two is blurrier than it looks. Weight updates are not as sweeping as people assume. Reinforcement learning turns out to touch only 5 to 30 percent of a model's parameters, and the same small subnetwork shows up again and again across random seeds Does reinforcement learning update only a small fraction of parameters?. So even the slow route behaves like a targeted adjustment rather than a rewrite.
At the other end, a lot of improvement happens with the weights frozen. Execution harnesses, meaning the code and procedures wrapped around a model, raised several models' scores on a hard terminal-use benchmark, and the same setup carried over to newer models unchanged Can execution harnesses lift model performance without retuning weights?. In another study, a stronger model wrote harnesses that nearly doubled a weaker model's Theory-of-Mind scores. It did this mostly by moving shaky reasoning steps into deterministic code Can a stronger model lift a weaker one at test time without retraining?. Filtering reasoning traces step by step, so weak traces can be stopped early, is another fast lever that never touches training Does step-level confidence outperform global averaging for trace filtering?. Between the two extremes sits representation finetuning. It learns small edits to the model's hidden activity while the weights stay frozen, and it is 10 to 50 times more parameter-efficient than LoRA Can editing hidden representations beat weight updates for finetuning?.
A third option is looping inside the architecture: running the same layers several times instead of adding more of them. Looped models gain reasoning ability through repeated depth rather than size Can models learn by looping instead of growing larger?. Looped world models report up to 100x parameter efficiency by spending extra passes on harder prediction steps Can looped computation replace parameter count in world models?. Looped mixture-of-experts models beat standard Transformers trained with the same compute Can looped models beat parameter-matched standard Transformers?. The same idea shows up in hardware: on phones, running a block twice is cheaper than loading a second block's weights Does recomputing weights cost less than moving them on mobile?. More looping is not always better, though. One coding model peaks at two passes, and from the third pass on the model oscillates instead of refining Does adding more loops always improve looped language models?.
The surprising part is why outside loops work at all. A standard LLM is bad at looping on its own. It can't run iterative numerical methods internally, so it pattern-matches to a memorized answer that looks plausible Do large language models actually perform iterative optimization?. It also can't take back a token once it has written it, which is exactly what solving constraint puzzles requires Why does autoregressive generation fail at constraint satisfaction?. So the clearest distinction is this: fast non-parametric loops supply abilities the architecture lacks (iteration, retraction, deterministic checking), while slow weight updates sharpen what the architecture can already do. The two are complements, not rivals.
Sources 12 notes
Across seven RL algorithms and ten LLM families, RL induces intrinsic parameter sparsity of 5–30% without explicit regularization. Critically, these sparse updates are nearly full-rank and nearly identical across random seeds, indicating structural rather than arbitrary parameter selection.
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Local step-level confidence catches reasoning breakdowns that global averaging masks and enables early stopping before traces complete. This approach achieves comparable accuracy gains to naive majority voting with far fewer generated traces, proving trace quality matters more than quantity.
ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.
Show all 12 sources
Models that re-apply layers in recurrent depth outperform larger feedforward networks on reasoning tasks. This works because recursion enables state tracking and compositional generalization that parameter scaling alone cannot achieve, with convergence signals providing natural halting.
LoopWM achieves up to 100x parameter efficiency by refining latent environment states through iterative computation in a shared block, with spectral-norm constraints providing formal stability guarantees. The approach mirrors physical system recurrence, spending more depth on harder prediction steps.
Loopie, a pair of MoE models with layer-loop recurrence, reportedly beats standard Transformers trained on the same pre-training compute budget. The win depends on co-designing recurrence with hardware-aware scaling and training efficiency, not recurrence alone.
MobileLLM shows that on memory-bound mobile hardware, sharing weights between adjacent transformer blocks by recomputing one block twice uses less latency than fetching separate weights, gaining accuracy with no parameter increase.
LoopCoder-v2 shows that two loops deliver broad gains over baseline, but three or more loops regress. Loop 2 carries the productive refinement; later loops oscillate with reduced representational diversity rather than converging toward better performance.
Research shows LLMs cannot perform iterative procedures in latent space. They recognize optimization problems as template-similar and emit plausible-looking but incorrect values, a failure mode that persists across model scale and training approaches.
The performance ceiling on constraint satisfaction problems is not a model-quality issue but an architectural limitation: autoregressive transformers cannot retract emitted tokens, while CSP solvers fundamentally depend on discarding invalid partial assignments. Symbolic solver integration works because it supplies what the architecture lacks.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- A Mechanistic Analysis of Looped Reasoning Language Models
- Loop the Loopies!
- Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
- LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
- Sharpening Tax in Post-Training
- Scaling Latent Reasoning via Looped Language Models
- Looped World Models
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach