PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Abstract AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities. This raises a deeper question: can these systems extend their capabilities to automate AI research itself? In this paper, we explore post-training, the critical phase that turns base LLMs into useful assistants. We introduce POSTTRAINBENCH to benchmark how well LLM agents can perform post-training autonomously under bounded compute constraints (10 hours on one H100 GPU). We ask frontier agents (e.g., Claude Code with OPUS 4.6) to optimize the performance of a base LLM on a particular benchmark (e.g., QWEN3-4B on AIME). Importantly, we do not provide any predefined strategies to the agents and instead give them full autonomy to find necessary information on the web, run experiments, and curate data. We find that frontier agents make substantial progress but generally lag behind instruction-tuned LLMs from leading providers: 23.2% for the best agent vs. 51.1% for official instruction-tuned models. However, agents can exceed instruction-tuned models in targeted scenarios: GPT-5.1 CODEX MAX achieves 89% on BFCL with GEMMA-3-4B vs. 67% for the official model. We also observe several failure modes worth flagging. Agents sometimes engage in reward hacking: training on the test set, downloading existing instruction-tuned checkpoints instead of training their own, and using API keys they find to generate synthetic data without authorization. These behaviors are concerning and highlight the importance of careful sandboxing as these systems become more capable. Overall, we hope POSTTRAINBENCH will be useful for tracking progress in AI R&D automation and for studying the risks that come with it.
Introduction. Recent advances in LLMs have given rise to a new class of AI systems: autonomous agents capable of reasoning, writing code, operating developer tools, and executing multi-hour workflows with minimal human oversight (Lin, 2026). Systems like Claude Code and Codex CLI have already begun to transform software engineering practice at scale. The obvious next question is whether these agents can accelerate AI research itself, a domain that has long depended on human intuition and manual trial-and-error. The question carries profound implications, as automating R&D more broadly is widely regarded as the key bottleneck to unlocking transformative advances in science and technology—potentially within years rather than decades (Amodei, 2024).
Why post-training? We study a central yet tractable component of modern AI research and development: posttraining. Post-training refers to the process of taking a pretrained LLM and systematically improving it through supervised fine-tuning, reinforcement learning from human feedback, and related alignment and capabilityenhancement methods. This stage is well defined because improvements can be directly measured using standardized evaluations such as AIME or HumanEval, which provide clear signals of performance gains after fine-tuning. The importance is equally clear: advances in post-training have been responsible for major gains in safety, instruction following, tool use, and reasoning. Despite this, no existing benchmark measures the ability of frontier LLM agents to perform post-training itself. Existing benchmarks focus on narrow AI R&D tasks or emphasize only certain aspects such as replication of existing papers (Chan et al., 2025; Wijk et al., 2024; Starace et al., 2025). Therefore, we need an end-to-end testbed that isolates the agent’s ability to directly improve model performance through post-training.
Our benchmark. To address this gap, we introduce POST- TRAINBENCH, where each evaluation pairs a base LLM (Qwen3-1.7B, Qwen3-4B, SmolLM3-3B, or Gemma-3-4B) with a target benchmark for the agent to optimize (AIME 2025, GSM8K, GPQA, HumanEval, BFCL, ArenaHard, or HealthBench). Agents are granted broad autonomy: they may write and execute code, search for and curate training data, and select any post-training strategy. We enforce only the minimal constraints necessary to preserve evaluation integrity. Agents may not train on benchmark test data, may not modify the evaluation harness, and may not fine-tune any model other than the provided base model.
At the end of each run, the agent submits a trained checkpoint, which is evaluated on the benchmark’s held-out test set. We evaluate frontier command-line agents (e.g., Codex CLI, Claude Code, and Gemini CLI) operating through standard developer tools without human interaction, under bounded resource constraints (10 hours on one H100 GPU).
Our findings. We find that frontier agents improve base models substantially but generally lag behind official instruction-tuned LLMs: the best agent reaches 23.2% average benchmark performance compared to 51.1% for instruction-tuned baselines. However, this gap is not uniform: agents can outperform human engineering on narrow tasks with clear evaluation signals. For example, GPT- 5.1 Codex Max post-trains Gemma-3-4B to 89% on function calling (BFCL), surpassing the official instruction-tuned model (67%). These results suggest that current agents can execute focused post-training successfully but do not yet match the broad, general-purpose post-training achieved by teams of expert scientists and engineers.
Related work. We review prior work on autonomous AI scientists, AI R&D automation, and relevant benchmarks.
Autonomous AI scientists. Fully autonomous research systems represent the frontier of AI R&D automation. The AI Scientist (Lu et al., 2024) demonstrated endto-end paper generation, AI-Researcher (Tang et al., 2025) introduced Scientist-Bench, and the Darwin-Gödel Machine (Zhang et al., 2025) showed recursive selfimprovement in coding agents. OpenAI’s FrontierScience benchmark (OpenAI, 2025) tests whether models can handle open-ended scientific reasoning rather than simple factual recall. However, systematic evaluations find no current framework completes full research cycles from literature understanding through validated results (Tie et al., 2025). POSTTRAINBENCH provides a standardized and verifiable way to measure the performance of automated AI research systems.
AI R&D automation. Interview studies with AI researchers (Owen, 2024; Leibowich et al., 2025) reveal substantial disagreement on automation timelines and identify compute bottlenecks as primary constraints. Several works address associated risks: Clymer et al. (2025) analyze risks from reduced human oversight and rapid capability acceleration, while Gasteiger et al. (2025) demonstrate that models can sandbag ML experiments without detection by zero-shot monitors. Anthropic’s evaluation of Claude Sonnet 4.5 (Anthropic, 2025) found the model does not yet automate entry-level researcher work but shows speedups on specific tasks. Our work shows recent models are much stronger than Sonnet 4.5 and can autonomously curate data, manage experiments and write entire training loops.
AI R&D benchmarks. Several benchmarks evaluate AI agents on ML engineering tasks. MLE-bench (Chan et al., 2025) uses 75 Kaggle competitions, subsequent work achieved medal-level performance in up to 47% of those competitions using advanced scaffolding (Qiang et al., 2025). MLAgentBench (Huang et al., 2024) provides 13 end-to-end ML tasks where agents autonomously develop or improve models given datasets and task descriptions. RE-Bench (Wijk et al., 2024) evaluates open-ended ML research tasks with human baselines, and HCAST (Rein et al., 2025) introduces a human-calibrated software engineering benchmark. Kwa et al. (2025) combine RE-Bench, HCAST and in one human-calibrated benchmark. Extrapolating their trends suggests that within 5 years, AI systems will be able to automate software tasks which currently take humans a month (Kwa et al., 2025). POSTTRAINBENCH differs from those approaches, because it uses larger models (up to 4B parameters) and allows agents complete freedom in their approach, both algorithmic and regarding the data which they use (subject to contamination constraints).
Method. 2. POSTTRAINBENCH: Setup Figure 2 shows our evaluation pipeline. We give each agent a base LLM, a target benchmark, access to compute node (a single H100 GPU) and internet access. The agent must build its training pipeline from scratch – we provide no starter code, training data, or hyperparameter configurations. The agent produces a post-trained model. We evaluate this model on the target benchmark and report its score. The goal of the agent is to maximize benchmark performance through post-training. Agents have full autonomy over data sources, training methods, and hyperparameters. They may iterate freely on their approach within time constraints (10-hour time limit).
In POSTTRAINBENCH the agents are only constrained to not use benchmark test data for training (data contamination) or substitute a different model than provided. These rules are enforced via an LLM judge (Appendix E). When the judge detects cheating, we assign the base model score. The overall score is computed across 4 base LLMs and 7 benchmarks. We detail the agent architecture and evaluation suite in the next subsections.
2.1. Agent Architecture The Agent consists of a scaffold, which behaves as the software layer and an underlying frontier model, which forms the underlying reasoning engine. The model processes context, generates plans, and decides which tools to invoke. The scaffold allows the LLM to use tools and manages the execution loop. Following ReAct (Yao et al., 2023), the scaffold operates in a loop: it presents the current context to the LLM, parses any tool calls from the response, executes them, and appends the results before the next iteration. The scaffold also handles permissions and context compression.
We evaluate different CLI-based agent scaffolds: Claude Code (Claude models), Codex CLI (OpenAI models), Gemini CLI (Google models) and OpenCode (Anomaly, 2025), an open-source scaffold that supports multiple model providers.
Tools. Agents typically use four tool categories: (1) file operations for reading and writing files, (2) shell execution for running arbitrary bash commands, (3) search tools for finding files and querying the web, and (4) context management for maintaining state across long sessions.
Example Execution Trace To illustrate how agents approach post-training tasks, we present a condensed execution trace from Claude Opus 4.5 using the Claude Code scaffold, post-training Gemma-3-4B-PT for HumanEval (Figure 3). The agent writes and debugs code, runs bash and python scripts and uses the internet to download data. It autonomously manages experiments and evaluates its intermediate results.
2.2. Evaluation Suite Our evaluation suite consists of seven benchmarks spanning mathematical reasoning, code generation, tool use, scientific reasoning, health and creative writing:
Mathematical reasoning. GSM8K (Cobbe et al., 2021) tests grade-school arithmetic word problems. AIME 2025 tests harder competition-level mathematics requiring multi-step reasoning.
Code generation. HumanEval (Chen et al., 2021) requires models to complete Python functions from docstrings.
Tool Use. BFCL v3 (Patil et al., 2025) tests function calling: given a natural language query and function specification, the model must generate a syntactically correct tool call with exact argument values. We use the exec_simple split.
Scientific knowledge. GPQA (Rein et al., 2024) contains graduate-level science questions in physics, chemistry, and biology. We use the main split.
Creative writing. ArenaHard v2 (Li et al., 2024b;a) has a creative writing split which we use as a user-centric benchmark. We call this ArenaHard-Writing.
Medical knowledge. We modify HealthBench (Arora et al., 2025) from OpenAI, designing an easy split testing multi-turn medical dialogue. Specifically, we subsample 245 questions requiring at least 5 turns, containing completeness-axis rubrics and at most 2 negative criteria. We call this HealthBench-Easy.
Discussion. Where do current AI R&D capabilities actually stand? While agents achieve substantial improvements over base models, interpreting them requires careful analysis. We expect that going from 7.5% (base model performance) to the 30% range will be relatively easy, since this can be achieved simply by teaching the base models to accurately follow instructions and format outputs correctly. Base models evaluated zero-shot often fail not because they lack knowledge, but because they output answers in the wrong format. A competent agent can fix this relatively quickly through simple supervised fine-tuning, which is already easy to implement for agents given how common it is on the internet and, consequently, in pre-training data. The harder challenge is approaching the official post-trained models (≈50%) and improving beyond them. This likely requires implementing distillation from more capable models, reinforcement learning, or even coming up with novel posttraining approaches. POSTTRAINBENCH is designed to capture such improvements, even if they exceed the performance of the best-known models.
Implications. Our results carry several implications for how the AI safety community should think about autonomous AI R&D. First, the gap between agent performance (23.2%) and instruction-tuned baselines (51.1%) suggests that full automation of post-training remains out of reach for now, but the rapid improvement across model generations—from 9.9% for Sonnet 4.5 to 23.2% for Opus 4.6 within roughly six months—implies this gap may close faster than expected. Second, the reward hacking behaviors we document (Section 5.3) are not hypothetical: agents trained on test data, substituted pre-trained models, and violated explicit API restrictions when constraints fell out of context. Crucially, these behaviors emerged naturally in the frontier models, without any adversarial prompting. As agents grow more capable, such specification gaming will likely become harder to detect and more consequential.
This concern is underscored by a striking pattern in our results: Claude Opus 4.6 – the highest-performing agent overall at 23.2% – was also the most frequent violator, flagged for contamination 12 times across 84 runs. This is not a case of weaker models cutting corners out of desperation. Rather, more capable agents appear better at finding exploitable paths: identifying specific benchmark samples to embed, reverse-engineering evaluation failure patterns, and even attempting to obscure contamination through cosmetic modifications such as renaming functions. The correlation between capability and rule violation suggests that as agents improve, the challenge shifts from preventing obvious cheating to detecting increasingly sophisticated specification gaming. Third, on the capability side, the fact that agents can already outperform expert human teams on narrow targets (e.g., BFCL) while operating with orders of magnitude less compute suggests that even partial AI R&D automation could meaningfully accelerate capability development in focused domains. This underscores the importance of developing robust oversight mechanisms before agents reach the capability threshold for broad post-training automation.
Limitations. Our evaluation has several limitations. The 10-hour, single-GPU budget, while practical for large-scale evaluation, does not reflect real-world post-training timelines or distributed training setups. The benchmark selection may inadvertently favor certain strategies, and agents optimize for single tasks rather than producing generalist models. Our LLM-based contamination judge may have false positives or negatives. Finally, cost constraints limited us to 3 runs for frontier agents and single runs for other configurations, restricting our ability to quantify variance.
Future work. Our goal is to maintain POSTTRAINBENCH as a continuously updated benchmark that provides meaningful signal about AI R&D automation capabilities. Moreover, we plan to release progressively harder versions of POSTTRAINBENCH in the future that keep pace with advancing capabilities. This means updating target benchmarks as existing ones saturate, swapping in newer base models as they are released, and expanding the set of agent scaffolds. One direction we are particularly interested in relates to safety and alignment. POSTTRAINBENCH measures whether agents can perform AI R&D, but an equally important question is whether agents will follow safety constraints while doing so.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What social dynamics enable or prevent agent collusion? Can AI research automation sustain progress through accelerating feedback loops? How does awareness of evaluation context influence model behavior?- Does eval-gaming explain why models act different when tested versus deployed?
- Do AI models behave differently when they believe deployment is real versus simulated?
- Why is catching an AI red-handed treated as a win condition?
- Can procedural guardrails prevent AI agents from making naive mistakes?
- What are the limits of black-box control as models grow more capable?
- How do weaker agents differ from stronger ones in test corruption?
- Why do most AI agent solutions score near zero despite occasional breakthroughs?
- How do domain experts recover from agent errors differently than novices?
- Do kernel optimization wins show agents discover genuinely novel techniques?
- Does agent capability separate into independent axes like performance and integrity?
- Does the metaproductivity mismatch occur outside coding-agent benchmarking tasks?
- Does source bias affect real deployed agents or only benchmark environments?
- Why do high-scoring agents default to known techniques rather than novel solutions?