What predicts success in ultra-long-horizon agent tasks?
Does an agent's initial solution quality matter more than its willingness to iterate? AUTOLAB's frontier-model benchmark suggests persistence through feedback loops may be the true differentiator.
AUTOLAB reframes what a long-horizon agent benchmark should test. Most agentic evals score either single-turn responses or short interactive trajectories; AUTOLAB instead hands the agent a correct but deliberately suboptimal baseline across 36 expert-curated tasks (system optimization, CUDA kernels, model development, puzzles) and asks it to improve the artifact within a strict wall-clock budget. The striking empirical result, across 17 frontier models, is that the dominant predictor of success is not the quality of the agent's initial attempt but its persistence — its willingness to repeatedly benchmark, edit, and incorporate noisy empirical feedback over many cycles. Most models, including proprietary ones, either terminate prematurely or exhaust their budget with minimal progress; claude-opus-4.6 is called out as a strong exception.
This is a sharper, more operational claim than "agents should iterate." It says the binding constraint is a behavioral disposition toward sustained empirical grounding, and that disposition is unevenly distributed across models that look comparable on one-shot benchmarks. It grounds How should we measure agent system performance beyond task success? with a concrete trajectory-level predictor, and it sits naturally alongside Does raw token spending actually predict agent performance? — persistence only pays if each loop returns informative, retained feedback, otherwise it is budget-burning churn, not progress.
The mechanism cuts against itself, however. Do models fail worse when their own errors fill the context? implies that more loops mean more accumulated mistakes in context, which should degrade the very iteration AUTOLAB rewards. The reconciliation is probably that persistence pays only when paired with calibrated scoring that lets the agent see whether an edit actually helped — pure persistence without trustworthy feedback would amplify error. That is why the authors single out harness design as the promising lever: the harness, not the backbone alone, decides whether long horizons compound feedback or compound noise.
Inquiring lines that read this note 106
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do single-axis benchmarks accurately measure agent capability for real deployment?- Do trajectory quality metrics predict agent safety and user trust?
- Can single-axis benchmarks measure across all three agent capability layers?
- What trajectory-level metrics replace one-shot task success measurement?
- What trajectory-level metrics matter beyond one-shot task success?
- What agent evaluation dimensions beyond task success does a single number hide?
- Do success-only evaluations systematically overestimate real-world deployment readiness?
- Should agent evaluation include trajectory quality beyond final success?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- Why is the coupled human-agent environment the right unit of evaluation?
- Does episode-level cost become the decisive factor when comparing AI agents in production?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- How can decision quality be automatically extracted from agent trajectories?
- Do kernel optimization wins show agents discover genuinely novel techniques?
- How do fixed external benchmarks anchor self-improving agent systems?
- Why do high-scoring agents default to known techniques rather than novel solutions?
- Why has agent research prioritized policy over world model development?
- What ecosystem conditions must exist for agents to function as economic participants?
- Can world models simulate actionable possibilities instead of just predicting next states?
- What domains let world models predict execution outcomes accurately enough?
- What counts as a mature governance model for agentic AI systems?
- Can agent-authored skill libraries compound autonomy gains over time?
- How can agent data flywheels improve task quality iteratively?
- Can gradients extracted with agent-level supervision transfer across different benchmarks?
- How does cross-agent supervision expand the set of convergent initial conditions?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- What makes next-state signals from agent trajectories a reliable learning source?
- Should evaluations shift toward open-world messy tasks instead of contests?
- Can automated benchmarks accurately capture progress on real-world long-horizon tasks?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- Can expert-frontier exams discriminate frontier capability better than saturated benchmarks?
- How do lab-scale benchmark tasks differ from real frontier AI research?
- Can frontier AI models match expert human performance on specialized tasks?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- How does poor belief tracking cause agents to keep acting past the point of usefulness?
- Why do autonomous agents report success on failed actions?
- Why do long-horizon agents fail when their models can solve individual steps?
- What causes delays between wrong decisions and visible consequences in long tasks?
- What within-run behavioral dimensions reveal where long-horizon agents succeed or fail?
- Why do most AI agent solutions score near zero despite occasional breakthroughs?
- Can objective search escape the limitations of fixed-objective central planning?
- How do goal and environment choices mediate AI agent risk pathways?
- What role should environmental rewards play versus human-specified objectives?
- How do agents differ in caution versus persistence across low-information scenarios?
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- Does intentionally varying environment properties isolate causal effects on agent performance?
- Should feedback channels be excluded from the reward path in agent evaluations?
- How does effective feedback retention govern long-horizon agent reliability?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- Can simple intrinsic reward signals emerge as effective drivers of complex capability in agents?
- Can agents improve reliably without an external standard?
- Can agents design their own objective functions as part of learning?
- What feedback signals matter most during harness evolution search?
- What role does effective feedback compute play in agent harness scaling?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- How would a parametric self-improvement loop differ from a non-parametric one?
- What external signals make self-improvement loops bounded rather than circular?
- When do diminishing returns appear in repeated cycles of AI self-optimization?
- Can empirical validation sustain long-term optimization without becoming gamed?
- What drives the mismatch between general benchmark leadership and task-specific performance?
- Which harness dimensions most directly predict agent system reliability?
- How much realized agent capability comes from the harness versus the model?
- What structural features drive instrumental convergence across different agent goals?
- Can autonomous agents detect increasingly sophisticated specification gaming as they improve?
- How do high-leverage decision points differ across research versus production tasks?
- What human decisions remain necessary even in closed-loop AI research venues?
- Which workplace tasks remain hardest for AI agents to complete autonomously?
- What task characteristics determine whether delegation can succeed?
- Does the planning-execution split between humans and agents depend on policy?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?
- Can judgment and accountability substitute for raw model capability in labor markets?
- What makes compute allocation a verifiable lever for pacing frontier AI development?
- Which bottleneck in the R&D feedback loop is the weakest link today?
- How does feedback latency from physical experiments shape AI system autonomy in research?
- Which parameters drive the largest uncertainty in AI R&D automation dates?
- Why do multi-year trials create inherent limits that model intelligence cannot overcome?
- What empirical parameters determine whether current AI loops are self-sustaining?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
grounds: supplies persistence/feedback-incorporation as a concrete trajectory-quality predictor
-
Does raw token spending actually predict agent performance?
Standard measures of agent effort—tokens, tool calls, operations—may not capture what makes inference-time scaling work. This explores what actually drives performance gains when agents spend more compute.
extends: persistence converts to progress only when each loop yields informative, retained feedback
-
Do models fail worse when their own errors fill the context?
As a model's prior mistakes accumulate in context, does subsequent accuracy degrade predictably? And can scaling or architectural changes prevent this self-contamination effect?
contradicts/qualifies: more iteration accumulates errors in context, so persistence helps only with trustworthy scoring
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
Original note title
on ultra-long-horizon optimization the predictor of agent success is persistence in the feedback loop not the quality of the first attempt