SYNTHESIS NOTE
Topics›Evolution›this note

Do self-improving agents really split into two distinct loops?

Explores whether modern self-improving agents can be understood through a clean abstraction separating fast scaffold updates from slow model weight updates, and whether this framework actually explains the field's recent progress.

Synthesis note · 2026-07-17 · sourced from Evolution

A survey of modern self-improving agents offers a clean systems abstraction: an agent is a configuration coupling a foundation model with an operational scaffold — prompts, memory, tools, and control logic — and self-improvement is a self-induced update operator that obtains and commits updates to either the model parameters or the scaffold components. That single distinction organizes the whole field, because it separates what changes from how fast it can change. Foundation-model improvement is the parametric, slower loop (driven by intrinsic generative demonstrations, intrinsic evaluative feedback, or extrinsic exploratory experience). Scaffolding improvement is the non-parametric, faster loop (updating prompts, memory, tools, or the full scaffold).

The value of the framing is that it makes the design space legible rather than a pile of named systems. It also explains why so much recent progress lives in the fast loop: updating a memory store or a prompt is cheap and reversible, while updating weights is expensive and risky. This is the same two-timescale structure surfaced empirically in Can agents adapt without pausing service to users? — but generalized from one system to a taxonomy. Which means the survey's real contribution is a shift of the object of study: the goal is no longer the static agent but the mechanism driving its evolution — the update operator itself and the signals that trigger it. Any concrete self-improvement design can be located by asking which loop it modifies and what signal drives the commit.

Inquiring lines that read this note 48

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What fundamental constraints limit how effectively agents can improve themselves? How does harness optimization generalize across different model architectures and domains? How do agent-learned skills transfer and improve across different tasks? Can harness architecture and protocols provide agent reliability without model scaling? What reasoning architectures enable models to solve complex problems efficiently? How do pretraining biases affect reward signal effectiveness in RLVR? What should agent evaluation prioritize to reveal reliable behavior? How does misalignment propagate through agent communication networks? When do multi-agent systems outperform single frontier models? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How do surface patterns enable correct outputs but reduce robustness? Can self-generated feedback reliably guide model training without ground truth? What trajectory-level metrics beyond task success best evaluate agent performance? How do standardized protocols improve multi-agent coordination and reliability? Why can't prompting alone inject genuinely new knowledge into models?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 118 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

modern self-improving agents divide into a slow parametric loop that updates model weights and a fast non-parametric loop that updates the scaffold of prompts memory and tools