AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Paper · arXiv 2610.02163 · Published October 1, 2026
Context Engineering

Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2% and 5.0%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction. training tasks and use a judge to review its compaction decisions, the summaries it writes, and its actions after compaction. The judge replaces flawed outputs with corrected ones before they are executed in the environment, so each trajectory continues from the corrected decisions and shows how to act after compaction. We use these trajectories for supervised fine-tuning (SFT), and then apply outcome-based reinforcement learning (RL), which jointly optimizes coding and compaction with binary task success as the only reward signal.

We evaluate AutoCompact on SWE-bench Verified (Jimenez et al., 2024; OpenAI, 2024) and SWE- PolyBench Verified (Rashid et al., 2025), achieving pass rates of 39.6% and 24.5%, respectively, with improvements of 9.2% and 5.0% over the base model.

Introduction. Large language model (LLM) agents have demonstrated strong capabilities on repository-level software engineering tasks by inspecting code, searching the repository, editing files, and running tests (Jimenez et al., 2024; Yang et al., 2024; Wang et al., 2025). Resolving complex issues often requires investigating multiple candidate causes and iterating between implementation and validation, so trajectories accumulate code snippets, tool outputs, and intermediate findings. As the agent moves from one stage of a task to the next, much of this information, such as exploratory hypotheses, failed attempts, and detailed tool outputs, becomes stale. Keeping the full trajectory thus fills the context with outdated details, whereas the next stage needs only a compact working state: the conclusions reached so far, the relevant code and workspace status, and the remaining actions. Effective context management therefore requires an agent to decide when to compact, what information to preserve in the working state, and how to continue execution from it.

Existing approaches address only part of this problem. Length-triggered methods such as CompactionRL (Li et al., 2026b) compact only when the remaining context budget falls below a fixed threshold, which ties compaction to context length rather than task progress: stale exploration accumulates until the threshold is reached, and compaction may then occur in the middle of an unresolved stage whose evidence is still needed. Proactive methods instead let the agent decide when to compact, either through an inference-time rubric (Li et al., 2026a) or through fine-tuning on trajectories into which compression calls are inserted offline (Liu et al., 2026). However, deciding when to compact is not sufficient: the working state may omit or misrepresent key information, and even when it is correct, the agent may fail to follow it, revisiting completed exploration or ignoring the intended next action. Neither kind of proactive method addresses these failures: a rubric provides no training signal, and offline insertion keeps the original actions after each inserted call.

We introduce AutoCompact, which trains a coding agent to decide when to compact, what to preserve, and how to continue as part of its policy. To collect training data, we run the base agent on

Related work. Coding agents. Research on repository-level coding improves task-solving capabilities through both harness design and model training. Work on harness design organizes repository exploration, code editing, and validation through agent-computer interfaces, tool execution environments, and structured repair workflows (Yang et al., 2024; Wang et al., 2025; Xia et al., 2025). Work on model training improves coding and problem-solving capabilities through executable task environments, expanded task and trajectory datasets, and supervised or reinforcement learning (Pan et al., 2025; Yang et al., 2025b; Wei et al., 2025). AutoCompact combines the two as model-harness co-design: the harness exposes a compact() action, and the model learns to use it jointly with task execution rather than through a separate compression module.

Memory and context management. Long-horizon agents manage growing interaction histories through memory hierarchies (Packer et al., 2023), abstraction of raw observations (Zheng et al., 2024), summaries of accumulated observations and interactions (Kang et al., 2026; Wu et al., 2025; Liu et al., 2026), or compression of sub-trajectories and selected history spans (Sun et al., 2025; Ye et al., 2025; Gao et al., 2026). These works mainly design how retained context is represented. AutoCompact uses a simple representation, a single working-state summary alongside the original task and recent turns, and focuses on when to compact and how to continue afterward.

Learning context-management policies. Rubric- and guideline-based methods (Li et al., 2026a; Kang et al., 2026) steer compaction through instructions rather than training the agent’s own compaction behavior. Supervised methods learn from constructed trajectories (Liu et al., 2026; Ye et al., 2025; Gao et al., 2026), for example by inserting context-management actions into completed trajectories (Liu et al., 2026), which keeps the original actions after each insertion, or by synthesizing them, as in concurrent work SWE-MeM (Gao et al., 2026), where a stronger model decides when to compress a span of steps and writes the summary but generates no task actions. RL methods (Wu et al., 2025; Sun et al., 2025; Li et al., 2026b; Gao et al., 2026) optimize context management through rewards, in SWE-MeM with additional step-level masks for memory-management failures, but do not directly supervise the actions that follow compaction. Unlike these methods, AutoCompact corrects the agent’s own compaction decisions, summaries, and subsequent actions during data collection, executes the corrected outputs, and refines this behavior with outcome-based RL alone.

Method. AutoCompact equips a coding agent with a compact() action (Section 2.1) and trains the policy in two stages. We first collect on-policy trajectories with judge-guided corrections to compaction timing, the generated working state, and the actions immediately following compaction (Section 2.2). These corrected trajectories provide the training data for SFT. We then apply outcome-based RL to jointly optimize coding and compaction using final task success as the reward (Section 2.3).

We augment the agent with a compact() action that can be invoked proactively during execution. The agent otherwise operates normally by inspecting files, editing code, and running tests. When it determines that the current stage has been sufficiently resolved and the accumulated history can be summarized, it may invoke compact() before reaching the context limit. For example, the agent may compact after localizing the cause of an issue and before proceeding to implementation.

Compaction replaces earlier interaction history with a model-generated working-state summary headed by # Auto Context Summary, while retaining the original task unchanged. The summary is intended to preserve established findings, relevant code and workspace state, and remaining actions, while omitting unnecessary details of prior exploration. The agent then resumes execution from the resulting context, as illustrated in Figure 1.

This mechanism makes compaction a model decision rather than solely a response to contextwindow pressure. Its effectiveness therefore depends on three behaviors: choosing an appropriate compaction point, constructing an accurate working state, and continuing reliably from that state. We next describe how we supervise these behaviors through judge-guided on-policy data collection.

We supervise three types of decisions corresponding to the main failure modes of compaction. Trigger correction checks whether the current point is appropriate for compaction: the agent should compact when the current stage has been sufficiently resolved, but continue exploring when intermediate evidence is still needed. Working-state correction checks whether the generated # Auto Context Summary accurately preserves the information needed for subsequent execution, including relevant conclusions, code and workspace state, and remaining actions. Continuation correction checks whether the first actions after compaction follow the resulting working state, rather than revisiting completed exploration or ignoring the intended next step, such as re-running a search whose result is already recorded in the working-state summary.

Corrections are applied before execution, so each one shapes the rest of the trajectory. When the judge identifies an output unsatisfactory, it writes a corrected output, which the environment executes in place of the original; the policy then generates the following actions from this corrected history. This differs from post-hoc annotation, where corrections do not influence later states.

Finally, we fine-tune the base policy on the corrected trajectories with the standard next-token prediction objective. The resulting policy, AutoCompact-SFT, learns both ordinary coding behavior and the three compaction behaviors introduced above: selecting appropriate compaction points, constructing actionable working states, and reliably continuing execution afterward. SFT provides a behavioral initialization for proactive compaction, but it optimizes the locally corrected decisions rather than their eventual effect on task completion. We then further optimize the policy using outcome-based RL, which directly targets task success.

Starting from the SFT checkpoint, we further train the agent on SWE-Gym (Pan et al., 2025) with end-to-end multi-turn RL (Xue et al., 2026), using group relative policy optimization (GRPO) (Shao et al., 2024) and a fully asynchronous RL system. Each rollout receives a binary outcome reward based on whether the final patch passes the task tests.

Discussion. RL improves compaction behavior without compaction-specific rewards. Figure 4 compares compaction behavior before and after RL. While Base rarely invokes compact(), AutoCompact- SFT already uses it on 44.3% of tasks, indicating that supervised training establishes the basic behavior. RL extends its use to 58.5% of tasks while reducing both types of summary omission: keystate omissions fall from 3.1% to 0.2%, and next-action omissions fall from 8.2% to 2.2%. These statistics are computed on trajectories for all 500 SWE-bench Verified tasks using keyword-based screening, supplemented by random manual spot checks. The two trends are complementary: compaction is used across more tasks, yet fewer summaries are flagged for missing state or next-action information. This suggests that task-success rewards encourage the agent to preserve information needed to continue the task from the compacted state.

Compaction contributes to the gains of the trained agent. The gains of AutoCompact may reflect both better agent behavior learned during training and the use of compaction during inference. To assess the latter contribution, we compare normal execution with a Summary-ignored variant of the same trained checkpoint. In this variant, compact() calls are skipped: no summary is generated, and execution continues with the existing history. Both conditions retain all learned parameters, allowing us to test whether executing compaction provides a benefit beyond the improvements encoded in those parameters. Figure 3 (d) shows that normal execution achieves higher pass rates, with the largest advantage under tight inference budgets. This indicates that the observed gains cannot be attributed solely to better agent behavior learned during training: part of the benefit depends on actually executing compaction. The advantage is 19.9% at $0.10 and remains 1.9% at $4.00, so executing compaction improves cost efficiency without reducing final performance.

Preserving the relevant state and specifying a next action do not by themselves make a summary usable. The summary-content metrics in Figure 4 check both, yet a summary can satisfy them while proposing a next action incompatible with the state it records, violating what we call summary selfconsistency. Figure 5 illustrates this gap through paraphrased summaries from AutoCompact-SFT and AutoCompact on the same SWE-bench Verified task.

This contrast highlights summary self-consistency as a dimension of summary quality beyond information coverage. The difference may stem from the training objectives: SFT imitates demonstrated summaries, which does not ensure that the recorded state constrains the next action, whereas RL rewards a summary only through the success of the actions that follow it.

Conclusion. AutoCompact trains coding agents to manage context as part of task execution. Judge-corrected trajectories provide supervision for compaction decisions, summary construction, and the actions that follow compaction, followed by outcome-based reinforcement learning that jointly optimizes coding and compaction. Experiments on SWE-bench Verified and SWE-PolyBench Verified show improved task success across all evaluated inference budgets, and ignoring the compaction calls of the same trained model lowers pass rates, showing that the gains do not come from training alone. These findings highlight the value of training not only the construction of compacted states but also the agent’s behavior that follows from these states, which prior methods leave unsupervised.

Limitations. and future work. Due to resource constraints, RL training uses 32K-token sequences, shorter than the 256K context window used at evaluation. Nevertheless, the gains across all evaluated budgets in both the 256K and 16K settings indicate that this training setting suffices to learn useful proactive compaction. More broadly, AutoCompact is a form of model-harness co-design: the harness provides the compaction mechanism, while the model learns when to invoke it, what to preserve, and how to continue afterward. We study this co-design within a single scaffold. Widely used agent harnesses such as Codex and Claude Code currently compact the context automatically when it approaches the window limit, which is the length-triggered design discussed in Section 1. Our 16K results suggest that combining such mechanisms with learned proactive compaction could further improve these systems, which we leave for future work.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do autonomous agents misreport success on failed actions? Why does self-revision amplify confidence in wrong model answers? How do models learn from self-generated outputs without cascading failures? How do curriculum design and feedback approaches affect model learning? Can base models hide emergent misalignment through alignment training? What limits language model accuracy in evaluating ideas?