SYNTHESIS NOTE
Topics›Novel Architectures›this note

Is long-context bottleneck really about memory or compute?

Explores whether the challenge of handling long context windows stems from storage capacity limits or from the computational cost of transforming context into internal state. Understanding this distinction reshapes how we design language models.

Synthesis note · 2026-05-28 · sourced from Novel Architectures

The standard framing of the long-context problem is capacity: attention scales poorly with context length, the KV cache grows, and we run out of room. "Language Models Need Sleep" reframes it as a compute-allocation problem. When the context window fills, the model enters a "sleep" — it performs N offline recurrent passes over the accumulated context and updates the fast weights in its state-space-model blocks through a learned local rule, then clears the KV cache and resumes. The information that would be lost on eviction is not stored verbatim; it is transformed into internal state by spending compute.

This relocates the bottleneck. The question is not "how much can we hold?" but "how much compute do we spend converting recent context into persistent weights, and when?" The design shifts that compute to the sleep phase, preserving wake-time prediction latency. The empirical signature confirms it is a compute story: increasing sleep duration N improves performance, with the largest gains on examples that require deeper reasoning — more offline compute buys more capability on hard cases, exactly the test-time-scaling pattern moved to an offline window.

The reframe is significant because it dissolves the capacity ceiling rather than raising it. A capacity solution adds memory; a compute solution adds passes. This connects to the vault's emerging theme that when a model thinks is as designable as how much — since When should AI systems do their thinking?, shifting inference to idle windows is a third temporal position for compute, and the sleep-consolidation mechanism is its architectural realization inside the weights. It also relates to alternatives that attack the capacity framing differently — since Can neural memory modules scale language models beyond attention limits?, one can add a long-term memory module instead of consolidating into fast weights. Counterpoint: spending compute on consolidation is only a win if the offline budget is genuinely free; under continuous load with no idle time, the sleep cost competes with serving. Why it matters: it tells architects to budget consolidation compute rather than chase ever-larger context windows.

Inquiring lines that read this note 163

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What compositional reasoning failures limit large language models despite scale? How should items be represented and indexed in recommenders? Can compression size predict model complexity better than parameter count alone? Can memory architectures handle ultra-long context better than attention? What structural properties of attention create systematic model biases? Why does memory consolidation cause performance regression in continual learning? How should inference compute be allocated based on problem difficulty? What reasoning architectures enable models to solve complex problems efficiently? Why can't prompting alone inject genuinely new knowledge into models? Can inference-time compute effectively substitute for model scale? What determines appropriate intervention timing and manner for AI agents? Is language model reasoning authentic and what causes models to reason? Can intelligent routing over smaller models outperform scaling a single large model? Can diffusion models match autoregressive performance on language generation tasks? Do language models reason like humans or mimic surface patterns? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How should retrieval systems handle complex multi-step reasoning? How do prompting refinements mask underlying biases and model frequency patterns? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Why does adding new knowledge through fine-tuning degrade existing capabilities? What causes reasoning models to fail or wander off track? How effectively can language models perform reasoning, especially combined with symbolic methods? Why do embedding systems fail to capture task-relevant relationships? What makes distillation transfer some model capabilities while suppressing others? What is the relationship between thinking tokens and reasoning accuracy? Can prompt-based context override biases that were embedded during pretraining? Can parallel reasoning outperform sequential reasoning under fixed token budgets? What mechanisms preserve shared understanding in evolving conversations? How should designers communicate what AI systems truly are and can do? What trajectory-level metrics beyond task success best evaluate agent performance? How should systems decide whether to retrieve or reason alone? How should agents manage memory granularity to improve long-term performance? What role does sparsity play in model behavior and scaling decisions? Can self-generated feedback reliably guide model training without ground truth? How do capability benchmark scores systematically misrepresent true model abilities? When do semantic similarity approaches miss structural retrieval failures? Can reasoning scale in latent space without tokens? What types of diversity prevent reasoning systems from collapsing? Do reasoning benchmarks predict model performance in long-horizon workflows? When do multi-agent systems provide sufficient quality returns on token investment? How does reasoning length affect model performance across different tasks? What fundamental constraints limit how effectively agents can improve themselves? How do standardized protocols improve multi-agent coordination and reliability? Can harness architecture and protocols provide agent reliability without model scaling? Does encoded knowledge in language models actually influence their outputs? How do neural networks achieve compositional generalization at scale? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 115 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the long-context bottleneck is compute to transform evicted context into internal state not memory capacity