Will agent experience overtake human data for AI progress?
Can AI systems improve primarily through their own environmental interactions rather than supervised learning from human examples? This matters because it shapes what kinds of agents we can build and what knowledge they can acquire.
Silver and Sutton argue that gains from training on human data are running out, and that the next source of AI improvement must be data an agent generates from its own interaction with an environment. The excerpt states the limit plainly: in "key domains such as mathematics, coding, and science, the knowledge extracted from human data is rapidly approaching a limit," and "the pace of progress driven solely by supervised learning from human data is demonstrably slowing." The successor data "must be generated in a way that continually improves as the agent becomes stronger," because "any static procedure for synthetically generating data will quickly become outstripped." They expect experience to "become the dominant medium of improvement," and suggest the shift "may have already started," citing AlphaProof's International Mathematical Olympiad medal as the first example.
The argument has a diagnostic half and a prescriptive half. The diagnosis is that human-centric RL, from RLHF onward, "side-stepped the need for value functions," let strong human priors reduce "the reliance on exploration," and left agents unable to "go beyond existing human knowledge." The prescription is four departures: agents "will inhabit streams of experience," their actions and observations will be "richly grounded in the environment," their rewards will be "grounded in their experience of the environment, rather than coming from human prejudgement," and they "will plan and/or reason about experience, rather than reasoning solely in human terms." Lifetime adaptation is the illustration: a health agent that tracks sleep and activity "over many months," or a science agent that runs simulations and proposes experiments. The safety case uses the same mechanism. Because an experiential agent's reward "may itself be adapted through experience," a misaligned reward "can often be incrementally corrected over time by trial and error."
Against the nearest notes, the paper is the field-level version of a critique that Can agents learn beyond what their training data shows? makes for a single training run: imitation turns the agent into a passive consumer of curated coverage. Can agents learn from their own actions without external rewards? offers one concrete route to the self-generated data the paper asks for, but it learns without an external reward, while Silver and Sutton keep a reward, grounded in the environment. The paper also assumes that experience turns into better behavior. Why do LLM agents ignore condensed experience summaries? finds that this conversion is not automatic, since agents ignored condensed experience even when it was the only input they were given. The sharpest contrast is with Do frontier AI agents actually conduct novel research or just optimize?. The excerpt says new insights "lie beyond the current boundaries of human understanding," while that study finds frontier models mostly compose established techniques, with genuine novelty rare. The two do not refute each other, since the paper gives no timeline, but the study shows where the gap between optimization and discovery currently sits.
The excerpt does not establish the timing or the mechanism. It is a position piece: it states that human-data progress is "demonstrably slowing" but shows no measurement, and it cites the AlphaProof result without describing how that system was trained, the passage breaking off mid-sentence. Its health, education and science examples illustrate what experience-driven agents could do, not evidence that they do. The authors also name costs they do not resolve: experiential learning "will increase certain safety risks," and moving away from human data "may also make future AI systems harder to interpret." The excerpt offers no method for recovering that lost interpretability. The reasonable reading is that the argument for the direction is coherent and well motivated, while its pace and its consequences for human understanding remain open.
Inquiring lines that read this note 1
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI systems determine and balance multiple competing objectives?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents learn beyond what their training data shows?
Explores whether supervised fine-tuning on expert demonstrations creates a hard ceiling on agent competence, or whether agents can generalize to scenarios their curators never captured.
the same passivity critique, applied to one training run rather than the whole field
-
Can agents learn from their own actions without external rewards?
Explores whether future states produced by an agent's own decisions can serve as supervision signals, bridging the gap between passive imitation learning and reward-dependent reinforcement learning.
one concrete mechanism for self-generated data, but without the environmental reward the paper keeps
-
Why do LLM agents ignore condensed experience summaries?
LLM agents faithfully learn from raw experience but systematically disregard condensed summaries of the same experience. This study investigates whether the problem lies in how summaries are made, how models process them, or whether models simply don't need them.
qualifies the paper's assumption that experience automatically becomes better behavior
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
contrasts the paper's predicted discovery beyond human understanding with measured limits on novelty
-
Can human data steer self-play RL toward human-compatible behavior?
Self-play RL finds effective but alien equilibria incompatible with human coordination. Can a small amount of human demonstration data redirect learning toward conventions humans actually use?
qualifies: human demonstrations still help as a light anchor on self-play reward, at 2500x less data than imitation learning
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Welcome to the Era of Experience
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- Quo Vadis, World Modeling?
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- LIMI: Less is More for Agency
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- Artifacts as Memory Beyond the Agent Boundary
- Large Language Model Agents Are Not Always Faithful Self-Evolvers
Original note title
Silver and Sutton argue that human data is approaching a limit and experience will become the dominant medium of AI improvement