Line of inquiry
Inquiring lines›How do knowledge organization and…›What mechanisms enable neural syst…›this line of inquiry
Do accumulated memories help or hurt continual learning in models?
A broader line of inquiry — a family of 73 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 73
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do long-term memory modules outperform consolidation into fast weights?
- Does moving memory outside model weights avoid the limitations of in-weight retention?
- Can precomputed inferences be stored in memory modules between model interactions?
- Why does specializing to one task make future task learning harder?
- Can data pruning strategies exploit the finite nature of memorization capacity?
- Why does fine-tuning for continuous space cause catastrophic forgetting?
- Why do accumulated memory systems sometimes hurt continual learning?
- Why do accumulated memories hurt continual learning more than no memory?
- Can in-weight memorization scale beyond model parameter count limits?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- How do adaptive memory modules compare to feedback-based working memory for long context?
- How does in-weight memorization scale with model parameter count?
- What makes factual memorization less efficient than tool-based retrieval?
- Why do external memory consolidation systems fail worse than naive in-context learning on continual tasks?
- How do retention gates regularize forgetting across different sequence model architectures?
- Why does semantic deduplication reduce memorization in fine-tuned models?
- Why does in-weight memorization fail compared to tool-based fact access?
- Can continuum memory systems prevent catastrophic forgetting in neural networks?
- Does composing multiple continual learning mechanisms reduce forgetting more than single approaches?
- Can episodic memory alone enable learning without parameter updates?
- How do complementary learning systems explain the need for fast and slow consolidation?
- Do models with unfilled memorization capacity appear to generalize falsely?
- How does in-weights adaptation create spurious forgetting in models?
- Can native memory procedures acquired through training handle stale or incorrect cached information?
- Why do large language models outperform fine-tuned models once repeated items are removed?
- How can a forgetting policy preserve rare knowledge while preventing over-generalization?
- Is forgetting in language models reversible or permanent knowledge loss?
- Why do pretrained model priors reduce the usefulness of retrieved experience?
- When does training a memory model beat RAG or fine-tuning?
- How do newly learned facts become accessible after gradient updates?
- Why does attending to own latents work better than bolted-on external memory stores?
- What mechanism transfers explicit memories into parametric model weights?
- Can neural modules memorize surprising tokens as adaptive long-term memory?
- Can document repetition accidentally memorize sensitive information instead of learning?
- How does distributional shift toward rare inputs change memorization reliance?
- Why is extracting training data insufficient proof that models memorize?
- What determines whether accumulated state generalizes spuriously across continual learning domains?
- Can adaptive memory modules combine long-term filtering with short-term attention benefits?
- Do retrieval-augmented memory systems actually solve the compartmentalization problem?
- How do memorization and attention map onto different memory systems?
- What causes catastrophic forgetting during domain knowledge embedding?
- Can memory-based adaptation and gradient fine-tuning operate on complementary timescales?
- Can we unlearn memorized text by finetuning only high-gradient weights?
- Can zero-weight drift through external memory replace parameter plasticity entirely?
- What gets lost when we describe memory as retrieval?
- Can episodic and semantic memory improve long-horizon task reasoning?
- How does memorization capacity saturation trigger the grokking transition?
- Why does recency-based recall outperform semantic similarity for episodic memory?
- Why does persistent memory alone fail to create genuine position-holding in models?
- How do trained weights differ from a stored library or text?
- Do sample-level similarities between pretraining and downstream tasks explain the frequency effect?
- How much does memorization capacity limit a model's ability to learn new information?
- Why does storing past judgments in memory make current evaluations worse?
- Does grokking in modular arithmetic follow the same three-phase learning trajectory?
- How can memory shift from a passive datastore to an actively trained component?
- Why does a replay mechanism prevent reasoner skills from over-specializing?
- What are the distinct sources of catastrophic forgetting in sequential fine-tuning?
- Why does grokking reveal the shift from memorization to genuine understanding?
- What distinguishes data that generalizes broadly from task-specific memorization?
- Can pretraining-frequency signals alone prevent RAG systems from confabulating about common knowledge?
- How do retrieved memories differ from decision-context passages for prediction?
- How does KL regularization prevent both forgetting and adaptation loss?
- Why does recall on demand not predict whether memory surfaces during user interaction?
- Can offline recurrent passes replicate sleep-based memory consolidation in AI?
- What is the theoretical capacity limit before memorization saturates?
- Can the joint-training principle extend beyond memorization and generalization pairs?
- What makes representation interventions more efficient than weight perturbations for finetuning?
- How do out-of-distribution tests reveal that optimization learning is memorization?
- Can a memory module be swapped between different base models?
- How does dual-rate learning separate episodic and procedural memory in neural networks?
- How do the three grokking phases connect to memorization capacity limits?
- What makes knowledge seeding equivalent to hippocampal replay in the brain?
- What distinguishes memory retrieval failures from failures to act on retrieved memory?