SYNTHESIS NOTE
Topics›Evaluations›this note

What predicts success in ultra-long-horizon agent tasks?

Does an agent's initial solution quality matter more than its willingness to iterate? AUTOLAB's frontier-model benchmark suggests persistence through feedback loops may be the true differentiator.

Synthesis note · 2026-06-27 · sourced from Evaluations
How does test-time scaling work for individual research agents?

AUTOLAB reframes what a long-horizon agent benchmark should test. Most agentic evals score either single-turn responses or short interactive trajectories; AUTOLAB instead hands the agent a correct but deliberately suboptimal baseline across 36 expert-curated tasks (system optimization, CUDA kernels, model development, puzzles) and asks it to improve the artifact within a strict wall-clock budget. The striking empirical result, across 17 frontier models, is that the dominant predictor of success is not the quality of the agent's initial attempt but its persistence — its willingness to repeatedly benchmark, edit, and incorporate noisy empirical feedback over many cycles. Most models, including proprietary ones, either terminate prematurely or exhaust their budget with minimal progress; claude-opus-4.6 is called out as a strong exception.

This is a sharper, more operational claim than "agents should iterate." It says the binding constraint is a behavioral disposition toward sustained empirical grounding, and that disposition is unevenly distributed across models that look comparable on one-shot benchmarks. It grounds How should we measure agent system performance beyond task success? with a concrete trajectory-level predictor, and it sits naturally alongside Does raw token spending actually predict agent performance? — persistence only pays if each loop returns informative, retained feedback, otherwise it is budget-burning churn, not progress.

The mechanism cuts against itself, however. Do models fail worse when their own errors fill the context? implies that more loops mean more accumulated mistakes in context, which should degrade the very iteration AUTOLAB rewards. The reconciliation is probably that persistence pays only when paired with calibrated scoring that lets the agent see whether an edit actually helped — pure persistence without trustworthy feedback would amplify error. That is why the authors single out harness design as the promising lever: the harness, not the backbone alone, decides whether long horizons compound feedback or compound noise.

Inquiring lines that read this note 106

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do single-axis benchmarks accurately measure agent capability for real deployment? Should governance of agentic AI systems be runtime or design-time? Can code harness improvements rival direct model scaling for capability? Can AI agents improve their skills through accumulated experience and reuse? How should agents coordinate through shared persistent code artifacts? Can smaller specialized models match frontier models on key metrics? How do real-world evaluations reveal AI capabilities that benchmarks hide? What makes agent memory systems durable and reusable across sessions? Why do autonomous agents misreport success on failed actions? How do AI systems determine and balance multiple competing objectives? Can AI systems discover fundamental improvements to their own architectures? How do agents learn to distinguish valuable feedback from noise? How do reward signal properties affect model reasoning and safety? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? What limits recursive self-improvement in autonomous AI systems? What explains the gap between benchmark scores and true reasoning capability? How does awareness of evaluation context influence model behavior? What prevents LLMs from applying their reasoning knowledge to improve outputs? How much of agent capability comes from harness versus the model itself? Why do standard evaluation practices obscure safety-critical AI failures? What distinguishes genuine communicative competence from surface language performance? When do multi-agent systems improve over single frontier models? Should agents compress episodic memory or retain raw interaction histories? How can evaluations be made robust against model reward hacking? What human oversight must AI research systems have? How should we measure frontier AI models' cyber exploitation capabilities? How should humans and AI agents share control and decision-making? How can humans maintain effective oversight as AI systems scale? Does AI deployment reduce or exacerbate workplace inequality and income instability? Which reinforcement learning modifications most improve dialogue quality in language models? Can language models reliably simulate personas and predict behavior? How do curriculum design and feedback approaches affect model learning? Can AI research automation sustain progress through accelerating feedback loops? Does pretraining establish the ceiling for what reward learning can improve? Why do confident AI outputs mislead human trust calibration? What gaps exist between benchmark performance and real deployment outcomes? What governance mechanisms can effectively constrain widely deployed AI systems? Does AI-assisted research sacrifice exploration breadth for productivity gains? How do evaluation environment design choices affect AI security? How does decomposing tasks into separate stages affect reasoning quality and safety? How does AI adoption reshape collaboration patterns in knowledge work?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 134 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

on ultra-long-horizon optimization the predictor of agent success is persistence in the feedback loop not the quality of the first attempt