If AI can do research, does one stubborn human-only step cap all the gains, or does AI just reshape the work around it?
Does Amdahl's law or partial substitutability better model research task automation?
This explores which simple economic picture better describes what happens as AI automates parts of research: Amdahl's law, where the steps that can't be automated cap the total speedup no matter how fast the rest gets, or partial substitutability, where AI can stand in for human research effort to some degree and the work reorganizes around it.
This explores whether automating research is limited by bottlenecks (Amdahl's law: the part you can't automate sets the ceiling) or by how well AI can substitute for human effort (partial substitutability: AI replaces some inputs imperfectly, and the work reshapes around it). The corpus has no papers that test either economic model directly. It does have plenty of evidence on the assumption the two models disagree about: is research a fixed set of separable steps, or does automation keep redrawing where the steps begin and end?
Amdahl's law needs a stable fraction of work that resists automation. Several findings suggest that fraction doesn't stay put. Can agents learn reusable sub-task routines from past experience? shows agents pulling reusable routines out of past work and stacking them into bigger ones. The gains grow as tasks get less familiar, so the automatable share expands with experience. Can a stronger model lift a weaker one at test time without retraining? makes the same point from another direction. A stronger model nearly doubled a weaker model's performance mainly by turning shaky reasoning steps into deterministic code. That isn't speeding up a fixed step. It changes which steps exist at all, which is what substitutability predicts.
The gains also often come from restructuring the work rather than making any one step faster. Can agent harnesses be automatically optimized across many environments? found that automatically redesigning an agent's scaffolding cut token use almost in half, separately from any improvement in the model. Can algorithms control LLM reasoning better than LLMs alone? gets its gains by deciding what context each step sees. And Do AIDE2's improvements transfer to unseen tasks? shows an automated research system whose improvements carried over to unfamiliar fields such as weather forecasting, so the substitution isn't confined to one narrow task.
The strongest case for Amdahl comes from mathematics. Can autoformalization work on individual statements alone? argues that translating a single theorem into formal, machine-checkable form only looks automatable because existing libraries quietly do the hard part. The real work is building a coherent web of definitions and supporting results that can't be split into independent pieces. When the parts of a task are tightly linked like this, automation can't route around the hard part, and the bottleneck picture fits.
Putting it together: substitutability fits better wherever research can be broken into modular pieces and reorganized. Amdahl-style ceilings reappear wherever the work only makes sense as a connected whole. Verification is a likely candidate for one of those hard steps. Can behavioral training prove a model always complies? shows that some guarantees can't be established from observed behavior at all, so a checking step may remain no matter how much else is automated. Your answer to the question may depend less on which law is true than on how much of a given field looks like modular engineering and how much looks like building a theory.
Sources 7 notes
Agent Workflow Memory induces sub-task routines at finer granularity than full tasks, abstracts example-specific values, and compounds them hierarchically. This produces 24.6% relative gain on Mind2Web and 51.1% on WebArena, with larger gains as train-test gaps widen.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.
LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.
The paper reports that AIDE2's improvements transfer to four held-out benchmarks spanning machine learning, algorithm engineering, and physics-based weather forecasting—the last being outside the selection task distribution. This demonstrates transferable gains beyond overfitting to the selection set.
Show all 7 sources
Real formalization requires theory-level work: even one theorem needs a coherent web of axioms, definitions, and lemmas. Statement-level approaches only succeed by borrowing from prebuilt libraries like Mathlib, hiding the actual complexity involved.
Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
- SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
- Agent Workflow Memory
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
- Rethinking the Evaluation of Harness Evolution for Agents
- Reasoning with Large Language Models, a Survey