SYNTHESIS NOTE
Topics›Correct but Not Understood›this note

Will agent experience overtake human data for AI progress?

Can AI systems improve primarily through their own environmental interactions rather than supervised learning from human examples? This matters because it shapes what kinds of agents we can build and what knowledge they can acquire.

Synthesis note · 2026-10-06 · sourced from Correct but Not Understood

Silver and Sutton argue that gains from training on human data are running out, and that the next source of AI improvement must be data an agent generates from its own interaction with an environment. The excerpt states the limit plainly: in "key domains such as mathematics, coding, and science, the knowledge extracted from human data is rapidly approaching a limit," and "the pace of progress driven solely by supervised learning from human data is demonstrably slowing." The successor data "must be generated in a way that continually improves as the agent becomes stronger," because "any static procedure for synthetically generating data will quickly become outstripped." They expect experience to "become the dominant medium of improvement," and suggest the shift "may have already started," citing AlphaProof's International Mathematical Olympiad medal as the first example.

The argument has a diagnostic half and a prescriptive half. The diagnosis is that human-centric RL, from RLHF onward, "side-stepped the need for value functions," let strong human priors reduce "the reliance on exploration," and left agents unable to "go beyond existing human knowledge." The prescription is four departures: agents "will inhabit streams of experience," their actions and observations will be "richly grounded in the environment," their rewards will be "grounded in their experience of the environment, rather than coming from human prejudgement," and they "will plan and/or reason about experience, rather than reasoning solely in human terms." Lifetime adaptation is the illustration: a health agent that tracks sleep and activity "over many months," or a science agent that runs simulations and proposes experiments. The safety case uses the same mechanism. Because an experiential agent's reward "may itself be adapted through experience," a misaligned reward "can often be incrementally corrected over time by trial and error."

Against the nearest notes, the paper is the field-level version of a critique that Can agents learn beyond what their training data shows? makes for a single training run: imitation turns the agent into a passive consumer of curated coverage. Can agents learn from their own actions without external rewards? offers one concrete route to the self-generated data the paper asks for, but it learns without an external reward, while Silver and Sutton keep a reward, grounded in the environment. The paper also assumes that experience turns into better behavior. Why do LLM agents ignore condensed experience summaries? finds that this conversion is not automatic, since agents ignored condensed experience even when it was the only input they were given. The sharpest contrast is with Do frontier AI agents actually conduct novel research or just optimize?. The excerpt says new insights "lie beyond the current boundaries of human understanding," while that study finds frontier models mostly compose established techniques, with genuine novelty rare. The two do not refute each other, since the paper gives no timeline, but the study shows where the gap between optimization and discovery currently sits.

The excerpt does not establish the timing or the mechanism. It is a position piece: it states that human-data progress is "demonstrably slowing" but shows no measurement, and it cites the AlphaProof result without describing how that system was trained, the passage breaking off mid-sentence. Its health, education and science examples illustrate what experience-driven agents could do, not evidence that they do. The authors also name costs they do not resolve: experiential learning "will increase certain safety risks," and moving away from human data "may also make future AI systems harder to interpret." The excerpt offers no method for recovering that lost interpretability. The reasonable reading is that the argument for the direction is coherent and well motivated, while its pace and its consequences for human understanding remain open.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do AI systems determine and balance multiple competing objectives?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 134 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Silver and Sutton argue that human data is approaching a limit and experience will become the dominant medium of AI improvement