Can dense subtask grading reveal agent progress on ultra-long tasks?
When agent tasks stretch to hours and hundreds of episodes, does breaking them into fine-grained graded subtasks expose meaningful progress that binary pass-fail scoring completely erases? This matters because most agents fail the final outcome anyway.
Long-Horizon-Terminal-Bench takes the standard Terminal-Bench setup — a reference solution or simulation engine per task — but decomposes each of its 46 tasks into fine-grained deterministic subtasks, each graded against the environment. This produces dense intermediate rewards and partial credit, which matters because the tasks are genuinely long: averaging 239 episodes, 9.8M tokens, and 88.9 minutes per run. On workflows that long, a binary final-outcome verdict discards almost all the information. The strongest model (Grok 4.5) passes only 28.3% at R≥0.95 while the mean across 17 models is just 6.4% — a near-floor result that outcome-only scoring would render as undifferentiated failure.
The design argument is that reward density is not a nicety but a measurement necessity once horizons stretch past one-shot problem solving. When agents make "meaningful but incomplete progress," a pass/fail benchmark cannot distinguish a model that got 80% of the way from one that stalled immediately — both read as zero. This grounds the broader claim that How should we measure agent system performance beyond task success?: dense subtask grading is one concrete instrument for that shift. It also echoes the training-side finding that Can we reward reasoning steps without human annotation? — dense signals expose structure that outcome-only aggregation flattens, whether the goal is evaluation or optimization.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What should agent evaluation prioritize to reveal reliable behavior? How can infrastructure records verify actual agent behavior? What trajectory-level metrics beyond task success best evaluate agent performance?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
extends: dense subtask grading is a concrete instrument for trajectory-level evaluation
-
Can we reward reasoning steps without human annotation?
Existing RL for reasoning uses only final-answer rewards, causing models to produce wastefully long chains. Can information theory provide dense, automatic feedback for individual reasoning steps?
grounds: the same dense-vs-sparse argument on the training side
-
Do automated benchmarks hide what frontier AI systems can really do?
Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?
relates: both target long-horizon realism that automated pass/fail benchmarks distort
-
Does planting honeypots in real coding tasks detect actual agent hacking?
Hack-Verifiable Terminal Bench moves honeypot detection from games to real-world coding tasks. But does a constructed shortcut measure the hacks agents actually find when deployed, or only how they respond to planted opportunities?
relates: a second modification of the Terminal Bench family; this note's dense grading adds resolution to what a pass covers, and that one adds a probe for whether a pass was earned (excerpt-only, no HVTB results)
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- Agents' Last Exam
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- OpenClaw-RL: Train Any Agent Simply by Talking
- Aspire: Can Models Self-Evolve from Vague Goals?
- FrontierChallenge: Evaluating Scientific Workflow Completion
Original note title
dense subtask grading on long-horizon terminal tasks turns agent evaluation from a pass-fail verdict into a progress measurement