Piling up experience doesn't make an AI agent better by itself; something has to turn that history into lessons.
Do agents actually convert raw experience into better behavior automatically?
This explores whether AI agents get better just by doing things and accumulating history, or whether something has to deliberately turn that history into lessons, and what the corpus says that something is.
This explores whether agents get better simply by piling up experience, or whether something has to turn raw history into lessons. The corpus says the conversion is not automatic. Experience only becomes better behavior when a specific mechanism does the work. Usually that mechanism sits outside the model, not inside it. The baseline problem is clearest with agents that never get raw experience at all. Agents trained only on expert demonstrations can't learn from their own failures, so their competence tops out at whatever the people who built the dataset thought to include Can agents learn beyond what their training data shows?. Interaction is necessary. It is not enough on its own.
What turns interaction into improvement? Look at Reflexion Can agents learn from failure without updating their weights?. After each attempt, the agent writes a short self-diagnosis in plain language and stores it for the next try. It works only because the environment gives an unambiguous success-or-failure signal. A clear 'you failed' leaves the agent no room to rationalize, so the reflection stays honest. Blurrier feedback lets an agent tell itself a flattering story about what went wrong. AgentFly Can agents learn continuously from experience without updating weights? takes the idea further. It splits memory into stores for past cases, subtasks and tools, so the agent can work out which past choices actually mattered and improve its policy without changing any model weights. Raw logs aren't lessons. Structure is what makes them usable.
The same goes for compression. If you summarize an agent's history carelessly, performance drops. DeepAgent Can agents compress their own memory without losing critical details? avoids this by folding past interactions into separate memories for events, current work and tools, and the folding step doubles as a moment for the agent to rethink its strategy. M3-Agent Can agents learn preferences by watching rather than asking? gets similar benefits by keeping 'what happened' separate from 'what I now know about this person.' That separation lets the agent infer someone's preferences from watching them, without having to ask. Across these papers, how memory is organized determines whether experience teaches anything or just piles up.
This is why much of the field now locates learning in the scaffolding around the model, often called the harness. A survey framework Do self-improving agents really split into two distinct loops? divides self-improvement into a slow loop, where model weights get retrained, and a fast loop that updates prompts, memory and tools. It finds that most recent progress is in the fast loop, because those changes are cheap and easy to undo. Reliability comes from moving memory, skills and protocols into the harness Where does agent reliability actually come from?. Harness improvements alone have lifted frozen models on hard terminal tasks Can execution harnesses lift model performance without retuning weights?. Automated search across many environments has even found harness changes that cut token traffic nearly in half without hurting performance Can agent harnesses be automatically optimized across many environments?. Experience shapes behavior here because engineers, and sometimes automated optimizers, build the channels it flows through.
The less obvious point is that real experience may not even be the bottleneck. Language world models trained to predict what an environment will do next can generate simulated experience, and agents trained on it outperformed agents trained in the real environments on three benchmarks Can language models learn to simulate agent environments?. One caution, though: experience only improves an agent on whatever it is measured against. Agents that win contest-style benchmarks still fail long, real professional workflows Why do agent benchmarks not predict real economic value?. So agents don't turn experience into better behavior on their own. They get better at whatever their feedback, memory structure and evaluation are built to reward.
Sources 11 notes
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Reflexion demonstrates that unambiguous environmental feedback (success/failure) enables agents to write useful self-diagnoses and improve across episodes without parameter updates. The binary signal prevents rationalization, and keeping reflections uncompressed preserves their usability.
AgentFly formalizes agent learning as a Memory-augmented MDP with three memory modules (case, subtask, tool) that enable credit assignment and policy improvement entirely through memory operations. The approach achieved 87.88% on GAIA validation without modifying LLM parameters.
DeepAgent's autonomous memory folding consolidates interaction history into episodic, working, and tool memory schemas. This reduces token overhead while letting agents pause to reconsider strategies—the autonomy and structure together avoid degradation that plagues poorly designed consolidation.
M3-Agent demonstrates that separating episodic events from semantic knowledge in an entity-centric graph, combined with parallel memorization and control processes, allows agents to infer and act on user preferences without asking. This architecture mirrors human cognitive systems that bind disparate information about individuals across sensory modalities.
Show all 11 sources
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
Research shows reliable LLM agents externalize three cognitive burdens—memory (state persistence), skills (procedural components), and protocols (structured interaction)—into a harness layer rather than relying on model scale alone. The harness unifies these externalities and eliminates the need for the model to solve the same problems repeatedly.
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.
Qwen-AgentWorld demonstrates that native language world models trained via next-state prediction on 10M+ trajectories outperform real-environment training on three benchmarks and transfer across seven domains, positioning next-state prediction as a foundation objective for agents.
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Useful Memories Become Faulty When Continuously Updated by LLMs
- Agent Learning via Early Experience
- Know It, Act on It: Investigating Memory Utilization in LLM Personalization
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- AgentFly: Fine-tuning LLM Agents without Fine-tuning LLMs
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- Rethinking Memory as Continuously Evolving Connectivity
- Survey on Evaluation of LLM-based Agents