SYNTHESIS NOTE
Topics›Data›this note

Can agents learn beyond what their training data shows?

Explores whether supervised fine-tuning on expert demonstrations creates a hard ceiling on agent competence, or whether agents can generalize to scenarios their curators never captured.

Synthesis note · 2026-05-03 · sourced from Data

The dominant paradigm for training language agents is supervised fine-tuning on expert-curated demonstrations. This bypasses the need for reward signals by letting agents map states to actions using static datasets. But the convenience hides a structural limitation: the agent never interacts with the environment during training, never observes the outcomes of its own actions, and therefore cannot learn from failure, refine its decision-making, or generalize to unseen situations.

The deeper problem is that the agent's competence is bounded by what the demonstration curators imagined. Every state-action pair in the dataset reflects a scenario someone thought to capture. Scenarios outside that imagination — edge cases, recovery from errors, paths the expert would never take — do not exist in the training signal at all. This means the agent learns the expert's idealized trajectory, not the structure of the environment. When the deployed environment presents anything unfamiliar, the agent has no internal model that can extrapolate, because its training never exposed it to consequences.

This is a passivity trap. Scaling high-quality human demonstrations is expensive and difficult to sustain, but even unlimited expert data would not solve the underlying problem — the agent is bound by the coverage of the demonstrations rather than by its own capacity to grow from experience. The demonstration paradigm assumes the world stops where the dataset stops.

The implication for agentic AI design is significant: data quantity and even data quality are insufficient. What agents need is the capacity to convert their own actions into learning signals — which is exactly what Can agents learn from their own actions without external rewards? proposes — requiring the agent to be in the environment, not merely trained on a snapshot of it.

Inquiring lines that read this note 168

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

When do multi-agent systems improve over single frontier models? Should governance of agentic AI systems be runtime or design-time? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can AI agents improve their skills through accumulated experience and reuse? How do agents learn to distinguish valuable feedback from noise? How do training data quality and composition affect downstream model performance? How can agents discover and adapt to user preferences during conversation? How should systems validate code that agents generate? What limits recursive self-improvement in autonomous AI systems? Can monitoring reasoning traces and behavior detect hidden agent deception? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? Can AI systems achieve real improvement without external human feedback? How effectively can test-time voting aggregate diverse reasoning samples? Does pretraining establish the ceiling for what reward learning can improve? Why does AI verification capability persistently exceed generation capability? How much of agent capability comes from harness versus the model itself? Does AI-assisted research sacrifice exploration breadth for productivity gains? Can language models reliably simulate personas and predict behavior? Which reinforcement learning modifications most improve dialogue quality in language models? What makes agent memory systems durable and reusable across sessions? Why do autonomous agents misreport success on failed actions? Can base models hide emergent misalignment through alignment training? Do accumulated memories help or hurt continual learning in models? Why do multi-agent systems reach premature consensus without genuine deliberation? Can artificial systems establish authority in domains requiring expert judgment? Can AI systems participate in genuine communication or only simulate it? Why do LLM research ideation systems generate novelty but lack diversity? Do single-axis benchmarks accurately measure agent capability for real deployment? What makes process supervision effective for training complex reasoning models? How do models learn from self-generated outputs without cascading failures? Can models develop genuine introspective capability, or only mimic it? Does AI assistance help or harm professional skill development? Why do language models struggle to implement user intent accurately from prompts? Should agents compress episodic memory or retain raw interaction histories? Does AI deployment reduce or exacerbate workplace inequality and income instability? Can AI systems discover fundamental improvements to their own architectures? How does awareness of evaluation context influence model behavior? How do multi-agent systems fail when coordination breaks down? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? What evaluation methods best detect reward hacking in AI agents? Why does polished AI output gain credibility despite fundamental verifiability problems? How do curriculum design and feedback approaches affect model learning? How do AI hiring systems affect authenticity, fairness, and candidate preferences? How do AI systems determine and balance multiple competing objectives? How can humans maintain effective oversight as AI systems scale? Do individually safe AI actions create unsafe outcomes in integrated systems? How should humans and AI agents share control and decision-making? How does AI adoption reshape collaboration patterns in knowledge work? How do real-world evaluations reveal AI capabilities that benchmarks hide?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 172 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

expert demonstrations lock agents into the imagination of the training data — restricting what an agent can learn to scenarios its curators happened to consider