SYNTHESIS NOTE
Topics›Evolution›this note

Does automated evolution match human-built agent performance?

Can an agent improved through automated loops in 8 days generalize as well as an agent refined through human-driven R&D? This tests whether autonomous design iteration reaches human-level quality on tasks outside the training set.

Synthesis note · 2026-09-24 · sourced from Evolution

The discussion's comparison: "On four held-out benchmarks spanning in- and out-of-distribution tasks, AIDE85 equals or surpasses AIDEhuman (section 3.3), a strong baseline developed through human-driven R&D (appendix B)." The evolved agent is labeled "AIDE85" in the excerpt, which does not define the label; I read it as the agent after the seven accepted rewrites, and that identification is mine.

The baseline is the point of the sentence. Seven rewrites accepted by an automated loop in 8 days are set against an agent that people improved by hand, and the paper's claim is that the loop's product is at least as good on tasks the loop did not select on. Read with Do AIDE2's improvements transfer to unseen tasks?, it says the automated route matched the human route on the held-out set.

What "equals or surpasses" does and does not say. It leaves room for ties on some benchmarks and wins on others, and the excerpt does not say which are which or by how much. It is a comparison to AIDEhuman, not to a test-time-search baseline at matched budget, so it does not stand in for the control in How should we measure gains from automatic harness evolution?. And it is a comparison with a human-driven R&D baseline, not with human–AI collaboration, so it does not test Can human-AI research teams improve faster than autonomous AI systems?: neither speed nor safety is compared in the excerpt, and the effort behind AIDEhuman (appendix B) is not given.

Inquiring lines that read this note 12

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does harness optimization generalize across different model architectures and domains? How can evolutionary algorithms maintain diversity during solution search? What makes imperfect LLM judges safe for optimization? Why does polished presentation create unearned authority in AI outputs? How does the generation-verification gap limit what we can measure about AI reasoning? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How do agent-learned skills transfer and improve across different tasks? What fundamental constraints limit how effectively agents can improve themselves? Can brute-force automated research substitute for iterative depth and human research intuition?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 84 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

on four held-out benchmarks the agent AIDE2 evolved equals or surpasses AIDEhuman, a strong baseline developed through human-driven R&D