Line of inquiry
Inquiring lines›How can multi-agent systems achiev…›What conditions allow multi-agent…›this line of inquiry
Can AI agents improve their skills through accumulated experience and reuse?
A broader line of inquiry — a family of 83 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 83
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do learned workflows transfer between different agents with minimal accuracy loss?
- Can skill repositories evolve toward execution-oriented refinement over time?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can agents learn beyond the boundaries of their training data curators?
- Do agents actually convert raw experience into better behavior automatically?
- Do weight-space skills lose detail compared to textual skill descriptions?
- How does real tool integration change what agents learn compared to simulated tools?
- Why do generic skill descriptions evolve into execution-oriented ones?
- Can skills learned from interaction trajectories outperform skills from static repositories?
- Why do trajectory-based skills fail to transfer across different environments and use cases?
- Can agents improve from deployment signals without explicit human annotation?
- What infrastructure decouples generation from training in asynchronous agent loops?
- What training method supports dynamic tool discovery in long-horizon agents?
- How can agent data flywheels improve task quality iteratively?
- Can context management policies transfer across agents of similar capability levels?
- Does a tight learning-rate bound prevent skills from escaping poor starting points?
- What makes next-state signals from agent trajectories a reliable learning source?
- Can agents teach each other skills without human supervision?
- How do agents automatically generate suitable learning tasks based on current capability?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- Can RL-trained meta-agents match or exceed manually designed workflows?
- Can next-state supervision work across different agent interaction types like conversations and tool calls?
- Can agents escape training data distributions without expensive real-world interaction?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- What domain properties determine whether causal rules transfer to new agents?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- Should production agents execute one tool or multiple tools per invocation?
- Can agents learn to use scaffolding structure the way they learn token weights?
- How can agents evolve their own skills without human input?
- How do task stream groupings provide long-horizon learning signals for curation decisions?
- Can tool adaptation work without freezing the agent in the loop?
- Why do imitation learning agents stay locked within human demonstration patterns?
- What makes an agent mechanism reusable versus benchmark-specific?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- Should optimal context budgets scale with agent competence or task complexity?
- Should we train the evolver or the executor when building self-improving agents?
- Can curriculum approaches teach agents when to stop exploring?
- How does cross-agent supervision expand the set of convergent initial conditions?
- How much does external context management transfer across similar capability agents?
- Can skill libraries prevent redundant narrow artifacts from proliferating?
- How do agents discover and select which tools to invoke?
- Does codifying domain rules into agent scaffolding work at library scale?
- How do agents decide which skills to chain together for a single task?
- Can simulation fidelity limit what agents learn from trained world models?
- Can graph topology represent successful trajectory clusters more effectively than skill libraries?
- How should AI skills be created and managed like software artifacts?
- Should agent capability be optimized separately from general capability?
- Can agents manage context through active delegation instead of progressive disclosure?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- Can influence estimation identify the most valuable trajectories in agentic training?
- How do fast and slow timescales enable continual agent adaptation?
- How much does agent performance depend on demonstration quantity versus curation quality?
- Can empirical testing replace semantic retrieval for finding useful skills?
- What happens when agents interact with environments and learn from their own mistakes?
- How do tool invocations drive agentic cost beyond token consumption?
- When should you optimize agent behavior versus tool performance separately?
- What prevents individual session learning from scaling system-wide?
- What lifecycle management prevents in-loop skill creation from bloating an agent?
- Why do current metacognitive training loops fail when agents encounter new domains?
- Why do environments authored once encode only what their builders imagined?
- What specific qualities make some demonstrations more effective for agency training?
- How do you prevent stale reward signals when skills evolve during deployment?
- What training difficulty and curriculum settings prevent instability in empathetic agent RL?
- Why does delegation training help models that work alone?
- How does mutual shaping through diverse training compare to population-level diversity effects?
- Can evolved algorithms transfer learning strategies across different datasets and tasks?
- What properties of agent systems only become visible across multiple sessions?
- How do agents retrieve and compose skills from hierarchical multimodal wikis?
- Why treat tutorial videos as a separate supply line from agent trajectories?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- Can tool-call advantage attribution distinguish between correct and incorrect calls in mixed trajectories?
- Can gradients extracted with agent-level supervision transfer across different benchmarks?
- Can agents acquire new skills online when offline skill coverage runs out?
- Can context management be optimized for an agent without retraining or changing the model?
- How do parametric and non-parametric updates differ in agents?
- Can curator modules trained on one executor transfer to entirely different agent backbones?
- How do agent capabilities change across 25 relay rounds of interaction?
- Does ontology investment actually improve agent performance on business tasks?
- Where does an agent's risk come from across its components and sequence?