Can agents learn new skills without forgetting old ones?
Explores whether externalized skill libraries—storing learned behaviors as retrievable code rather than parameter updates—can solve the catastrophic forgetting problem that plagues continual learning systems.
VOYAGER introduces an architecture for lifelong learning that solves the catastrophic forgetting problem through externalization rather than internal parameter management. Three components work together:
Automatic curriculum — proposes tasks based on the agent's current skill level and world state (finding yourself in a desert means harvesting sand before iron). Generated by GPT-4 with the overarching goal of "discovering as many diverse things as possible" — an in-context form of novelty search.
Ever-growing skill library — each successfully completed task produces an executable code program stored in the library, indexed by the embedding of its description. When similar situations arise, relevant skills are retrieved by semantic similarity. This externalizes learned behavior as retrievable artifacts rather than weight updates.
Iterative prompting with environment feedback — incorporates execution errors, environment feedback, and self-verification for program improvement. The agent refines skills based on real-world outcomes.
The compounding mechanism is the key insight: complex skills are synthesized by composing simpler programs. Fighting zombies builds on combat primitives; navigating a cave builds on movement and resource-gathering skills. This composition enables rapid capability growth without the forgetting that plagues weight-update-based continual learning methods.
Three lifelong learning requirements are met: (1) propose suitable tasks based on current capability and context, (2) refine skills from environmental feedback and commit to memory, (3) continually explore in a self-driven manner. These parallel the three requirements of the When should proactive agents push toward their goals versus accommodate users? framework — goal awareness, context adaptation, and initiative.
Because Can agents learn from failure without updating their weights?, VOYAGER's skill library is a more structured version of the same principle: externalize learning as retrievable artifacts. The embedding-indexed retrieval means skills are found by semantic similarity, not exact match — enabling transfer to novel but related situations.
Since Can communication pressure drive agents to learn shared abstractions?, the skill library pattern may generalize: agents under performance pressure naturally develop reusable, composable abstractions.
MUSE-Autoskill generalizes Voyager's compounding library into an explicit five-stage skill lifecycle — creation, memory, management, evaluation, refinement — turning skills from disposable generation outputs into "long-lived, experience-aware, testable assets." Two extensions matter for the catastrophic-forgetting claim. First, skills are validated through unit tests plus runtime feedback, so the library does not just grow but is continuously checked for reliability — addressing the gap where Voyager stores any successfully-executed program regardless of later robustness. Second, MUSE adds skill-level memory that accumulates per-skill experience across tasks, so reuse improves over time rather than staying static after first synthesis. On SkillsBench, generated skills reach 87.94% on their tasks and transfer to other agents with minimal accuracy loss, evidence that lifecycle management (not just synthesis) is what makes externalized skills durable infrastructure.
Inquiring lines that read this note 185
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does model capacity affect learning performance on diverse downstream tasks? What makes agent memory systems durable and reusable across sessions?- How should GUI agents remember patterns across different software environments?
- Can environmental scaffolding replace internal memory scaling in agent design?
- What memory and planning capabilities do AI companions need for evolving user needs?
- Why do memory and feedback loops matter more than model size for agent reliability?
- What makes memory trajectories topologically stable under persistent reuse?
- Can episodic memory of UI traces improve open-world agent adaptation?
- Can state-indexed memory retrieval breadth predict gains in web agent robustness?
- How does PRAXIS differ architecturally from Agent Workflow Memory and causal rule learning?
- Can topology repair fix consolidation failures in agent memory?
- Can pruning policies alone solve working memory bloat in agents?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
- Can AI models retain knowledge across changing environments without catastrophic forgetting?
- What distinguishes working memory from strategic memory in agent task execution?
- How does durable memory quality shape agent performance over time?
- What can agents learn from the brain's complementary learning systems?
- How does external context control compare to agents managing their own state internally?
- Can externalizing bookkeeping to a stateful harness replace internalized memory control?
- What causes multi-turn agent failures: weak memory control or missing knowledge?
- How should agent memory links evolve based on execution feedback?
- Can workflow memory compound reusable skills into measurable success improvements?
- Why do persistent AI systems require fundamentally different design than ad-hoc supporters?
- What discarding policy prevents both stale entries and loss of rare critical knowledge?
- What happens to agent performance when stored knowledge continuously updates?
- How do workflow and function memories contribute differently in agent learning?
- Can persistent memory architectures enable AI to reuse and stabilize invented concepts?
- How do modern agents separate fast non-parametric updates from slow weight learning?
- Which AI interaction patterns preserve learning while which ones degrade skill formation?
- Does outsourcing tasks to AI reduce opportunities for skill development?
- Does AI training preserve learning that transfers to independent subsequent tasks?
- Do AI productivity gains require existing skills or enable learning new ones?
- Why do skill-learning barriers prevent workers from adapting to AI tools?
- How does tacit knowledge spread when incomplete contracts cannot finance training?
- Why does persistent memory alone fail to create genuine position-holding in models?
- Why does fine-tuning for continuous space cause catastrophic forgetting?
- Can continuum memory systems prevent catastrophic forgetting in neural networks?
- Can episodic memory alone enable learning without parameter updates?
- How do retention gates regularize forgetting across different sequence model architectures?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- How can memory shift from a passive datastore to an actively trained component?
- Why does a replay mechanism prevent reasoner skills from over-specializing?
- Can offline recurrent passes replicate sleep-based memory consolidation in AI?
- Why does specializing to one task make future task learning harder?
- How does KL regularization prevent both forgetting and adaptation loss?
- Can zero-weight drift through external memory replace parameter plasticity entirely?
- How does in-weights adaptation create spurious forgetting in models?
- How can a forgetting policy preserve rare knowledge while preventing over-generalization?
- What determines whether accumulated state generalizes spuriously across continual learning domains?
- Why do accumulated memory systems sometimes hurt continual learning?
- Why do external memory consolidation systems fail worse than naive in-context learning on continual tasks?
- Can native memory procedures acquired through training handle stale or incorrect cached information?
- Does composing multiple continual learning mechanisms reduce forgetting more than single approaches?
- What are the distinct sources of catastrophic forgetting in sequential fine-tuning?
- Do dynamic environments enable different kinds of agent-environment coevolution?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- Can combinational creativity alone drive open-ended learning in agents?
- How does component-level self-evolution prevent information loss in multi-agent trajectories?
- Can small numbers of curated demonstrations produce emergent agentic behavior?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- Why does capability discovery become the bottleneck in large agent systems?
- How does deterministic feature engineering increase information for computationally bounded agents?
- Can stochastic memory movement converge to better team strategies?
- Can tool adaptation work without freezing the agent in the loop?
- How does real tool integration change what agents learn compared to simulated tools?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- What happens when agents interact with environments and learn from their own mistakes?
- What training difficulty and curriculum settings prevent instability in empathetic agent RL?
- Can agents improve from deployment signals without explicit human annotation?
- Can curriculum approaches teach agents when to stop exploring?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- Can curator modules trained on one executor transfer to entirely different agent backbones?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Do learned workflows transfer between different agents with minimal accuracy loss?
- How do agents automatically generate suitable learning tasks based on current capability?
- How do you prevent stale reward signals when skills evolve during deployment?
- Can skill libraries prevent redundant narrow artifacts from proliferating?
- What lifecycle management prevents in-loop skill creation from bloating an agent?
- What training method supports dynamic tool discovery in long-horizon agents?
- Why do current metacognitive training loops fail when agents encounter new domains?
- How do fast and slow timescales enable continual agent adaptation?
- What properties of agent systems only become visible across multiple sessions?
- Should we train the evolver or the executor when building self-improving agents?
- Can agent skills move from prompts to trainable parameters?
- How can agents evolve their own skills without human input?
- Can agent-authored skill libraries compound autonomy gains over time?
- Can agents learn to use scaffolding structure the way they learn token weights?
- How should AI skills be created and managed like software artifacts?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- Can agents teach each other skills without human supervision?
- Can skill repositories evolve toward execution-oriented refinement over time?
- How do agents decide which skills to chain together for a single task?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- How can agent data flywheels improve task quality iteratively?
- Can agents acquire new skills online when offline skill coverage runs out?
- How do agents retrieve and compose skills from hierarchical multimodal wikis?
- Why treat tutorial videos as a separate supply line from agent trajectories?
- What prevents individual session learning from scaling system-wide?
- What makes next-state signals from agent trajectories a reliable learning source?
- Can agents learn beyond the boundaries of their training data curators?
- Why do imitation learning agents stay locked within human demonstration patterns?
- What makes self-modifying architectures learn their own update rules?
- What role does self-learning play in improving agent reasoning without annotation?
- Can AI systems improve themselves without external feedback?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Can agents extract structured lessons from failure without massive compute budgets?
- Why do deployed models lack the learning and planning Weinstein attributes to them?
- Does narrow reallocation to remaining tasks constitute genuine adaptation?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?
- Can self-distillation reduce catastrophic forgetting in continual learning?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- How do evolutionary archives enable diverse exploration in self-improving systems?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- Can applicability conditions and veto rules make self-training stable across substrates?
- Do evolutionary archives let agents improve themselves without formal proof?
- How do evolutionary archives enable open-ended self-improvement without formal proofs?
- What alternatives exist when required knowledge is absent from training?
- How much can externalized skills improve models before hitting diminishing returns?
- How does scaffolding unstable mechanics improve reinforcement learning for search?
- What details do high-level trajectory abstractions lose that state-grounded recall preserves?
- When does memory consolidation help agents instead of hurting performance?
- Why do continuously consolidated agent memories eventually degrade below no-memory baseline?
- Why does memory consolidation degrade agent performance below baseline?
- How can agents distinguish over-generalized lessons from genuinely useful long-tail knowledge?
- Why does higher agent recall make forgetting problems harder?
- Why do agents systematically ignore condensed experience in their skill documents?
- Can agents learn from their own experience without fine-tuning through episodic memory?
- What shapes of memory help frozen agents improve without retraining?
- Does operator-conditioned memory let search compose learned behaviors more effectively?
- Does training on granular tasks beat training on the full function calling problem?
- Can models recover knowledge with completely unrelated retraining tasks?
- What knowledge injection routes trade flexibility against training cost?
- Why do completion-mode strengths not transfer to agentic settings?
- What makes idle window detection valuable for continuous agent improvement?
- How do agents discover and construct new APIs from existing applications?
- Can skill validation through testing prevent unreliable programs from accumulating?
- How do skills authored in-loop validate faster than offline generated skills?
- How should agents decide which created code is worth persisting?
- How should evolving systems track lineage and enable rollback of changed mechanisms?
- What execution-layer design prevents agents from passively reacting to environments?
- Why does externalized state beat parameter scaling for agent reliability?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- What makes agent-initiated artifacts the underexplored frontier in harness engineering?
- Why does the harness layer accumulate distributed behaviors over time?
- What makes behavior localization the bottleneck in agent harness evolution?
- Why do persistent, resynchronized artifacts compound harness capability gains?
- How should humans specify deterministic abstractions of RL problems?
- Can reinforcement learning add new capabilities or only remove inaccurate knowledge?
- What hard-to-verify tasks will remain resistant to reinforcement learning?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- What distinguishes learnable perturbations from bifurcation-triggered motive shifts?
- When should agent-created code be promoted into permanent harness infrastructure?
- How do agents decide which created code should persist versus disappear?
- How should human oversight apply to persistent agent-authored code?
- Can one-off agent code be safely promoted to durable infrastructure?
- Can versioned capability vectors solve the discovery gap in existing protocols?
- How do agents decide which created code deserves long-term persistence?
- Are durable shared code artifacts better than per-task harness patches?
- Can applicability conditions be preserved automatically when agents reflect on trials?
- What makes exploration and reflection rewards verifiable in agentic environments?
- How does SDPO relate to agents learning from verbal reflection without parameter updates?
- What mechanisms let later agents inherit information left by earlier ones?
- How does effective feedback retention govern long-horizon agent reliability?
- How can we reorganize repositories to make behaviors easier to locate?
- How does source-blind reconstruction verify that extracted skills are specific enough to be reusable?
- Why do evolved harness edits mostly memorize rather than generalize?
- Do evolved harness edits capture reusable strategies or task-specific memorization?
- Can harness evolution be redirected from memorization toward strategy distillation?
Related concepts in this collection 10
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents learn from failure without updating their weights?
Explores whether language models can improve through trial and error by storing reflections in episodic memory rather than fine-tuning. This matters because it suggests a fundamentally different path to agent adaptation.
related architecture: episodic memory as external learning
-
Can communication pressure drive agents to learn shared abstractions?
Under what conditions do AI agents develop compact, efficient shared languages? This explores whether cooperative task pressure—rather than explicit optimization—naturally drives abstraction formation, mirroring human collaborative communication.
same pattern: reusable abstractions under optimization pressure
-
When should proactive agents push toward their goals versus accommodate users?
Proactive dialogue agents face a tension between reaching their objectives efficiently and keeping users satisfied. This question explores whether these two aims can coexist or require constant negotiation.
parallel requirements for autonomous goal setting
-
Does self-generated training data improve model learning?
Can models learn more effectively from training data they generate themselves rather than data created by external sources? This explores whether a learner's own restructuring process produces better learning outcomes.
SEAL: model-specific data as capability building blocks
-
Can agents learn continuously from experience without updating weights?
This explores whether LLM agents can adapt to new tasks and failures by retrieving past experiences from memory alone, rather than requiring expensive parameter fine-tuning or rigid hardcoded rules.
AgentFly composes cases where VOYAGER composes skills; both achieve continual learning without parameter updates, but AgentFly adds a Q-function for principled case retrieval beyond static similarity
-
Can neural networks learn compositional skills without symbolic mechanisms?
Do neural networks need explicit symbolic architecture to compose learned concepts, or can scaling alone enable compositional generalization? This asks whether compositionality is an architectural feature or an emergent property of scale.
VOYAGER's skill library implements compositional generalization externally: complex skills are synthesized from simpler skill programs, achieving the linear-scaling efficiency the MLP proof demonstrates; the embedding-indexed retrieval ensures the training distribution covers the compositional space
-
Can we teach LLMs to form linguistic conventions in context?
Humans naturally shorten references as conversations progress, but LLMs don't adapt their language for efficiency even when they understand their partners do. Can training on coreference patterns teach this convention-forming behavior?
both VOYAGER and convention formation involve agents developing compact reusable abstractions through interaction: skills are behavioral conventions for task completion, and linguistic conventions are communicative skills for efficient reference; the shared mechanism is that repeated interaction under performance pressure drives abstraction
-
What happens to code that agents create and then share?
Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.
exemplifies: a compounding skill library is a concrete case of persistent agent-authored artifacts the frontier asks about
-
Does creating skills inside the agent loop eliminate mismatches?
Can coupling skill creation directly to the runtime reasoning loop—rather than authoring skills offline—close the gap between when skills are made and when they're used? This matters for whether agents can ground new capabilities in their actual situated context.
extends: Voyager builds the library by synthesis; MUSE specifies that creation happens in-loop where consumed, preventing creation-usage mismatch
-
Can frozen models learn better by extracting context into skills?
When a model encounters unfamiliar material in its context, can we help it reason more effectively by explicitly extracting rules and procedures from that material rather than changing the model itself?
grounds the accumulating store in a primitive: single-context skill extraction is the unit the compositional library scales and compounds
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Demystifying Agent Skills: Why They Work-Until They Don't
- MetaClaw: Just Talk — An Agent That Meta-Learns and Evolves in the Wild
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Original note title
compositional skill libraries that compound through synthesis enable lifelong learning without catastrophic forgetting