Sharpening Tax in Post-Training

Paper · arXiv 2610.01509 · Published October 1, 2026
Training and Fine-Tuning

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.

Introduction. Post-training at scale, particularly via reinforcement learning (RL), has transformed pre-trained general text completion machines into reasoning models (Shao et al., 2024; Guo et al., 2025). Beyond competition-level math and coding reasoners, post-trained large language models (LLMs) now serve as goal-oriented agents that call tools, manage long context, and interact with an environment across multiple turns (Anthropic, 2025; OpenAI, 2025; Google, 2025; Meta Superintelligence Labs, 2026; xAI, 2026). A common belief behind this progress is that frontier-level reasoning capability emerges during intensive post-training, e.g., RL unlocks new capabilities and pushes the reasoning boundary beyond what the pre-trained base model can already reach (Uesato et al., 2022; Guo et al., 2025). However, this belief has recently been called into question.

Does RL post-training create fundamentally new capabilities, or does it merely amplify a few rewarding behaviors that the base model already possesses, i.e., distribution sharpening? This question has attracted broad attention (Yue et al., 2025; Zhao et al., 2025; Wu et al., 2025; Yuan et al., 2026b; Wen et al., 2026; Shen et al., 2026b; Zhou, 2026). A dominant observation so far is that an RL-trained policy gains accuracy (pass@1) at the expense of solution coverage (pass@K) (Yue et al., 2025; Zhao et al., 2025). Despite lots of reasonable positions, from advocacy to skepticism, with training-dependent (Liu et al., 2025a) and data-dependent (Zhang et al., 2025; Shen et al., 2026a) viewpoints in between, current evidence is mainly limited to math and coding tasks (Yue et al., 2025; Zhao et al., 2025; He et al., 2025; Shao et al., 2026). Unfortunately, LLM coverage analysis in these domains does not settle the question since (1) pre-training provides enormous exposure to math and coding; (2) evaluation often checks only a final answer, so a lucky rollout with flawed reasoning earns credit (Skalse et al., 2022; Pan et al., 2024; Wen et al., 2026).

We argue that agentic tasks provide a stress testing evaluation suite to test the sharpening hypothesis. Unlike math and coding problems, for which a base model may already encode many candidate solution strategies, a model in agentic tasks must execute structured actions (Patil et al., 2025), invoke diverse tools through scenario-specific interfaces (Schick et al., 2023), incorporate heterogeneous environmental feedback under uncertainty (Oh et al., 2026b), and stay coherent over long-horizon trajectories (Froger et al., 2026). Those behaviors are far less common in pre-training data mix and are widely believed to be learned in post-training (Zeng et al., 2024; Su et al., 2026). Therefore, one might expect that post-training does more than redistribute probability mass among behaviors already present in the base model; it may be necessary to make successful agentic trajectories reachable at all.

To test this expectation, we systematically study LLMs’ reasoning boundaries on agentic tasks before and after post-training. In particular, we first focus on the publicly available open-source checkpoint pairs of the base and post-trained LLMs to understand how large-scale general post-training affects the agentic reasoning boundary of base LLMs. Our surprising observation (§3) is that while the post-trained LLMs lead in single-shot accuracy, pre-trained base LLMs can also perform agentic reasoning when paired with a lightweight harness, and they even surpass their post-trained counterparts in terms of solution coverage given enough test-time compute. We then analyze how post-training changes the per-task success probability distribution, revealing that it bimodalizes the mass towards two extremes, always solved or never solved, thereby forgoing the benefits of test-time parallel scaling of multiple rollouts to improve single-shot accuracy.

Motivated by these test-time scaling dynamics, we propose Sharpening Tax (§4), a diagnostic metric that quantifies how much post-training shrinks the effective reasoning coverage as a single scalar. Across 14 model backbones from popular open source model families and three representative basic benchmarks for agentic tasks, we find that the tax is pervasive, implying that modern post-training consistently trades coverage for sampling efficiency; then we highlight the practical usefulness of Sharpening Tax, which can be estimated from a handful of rollouts to predict the future tax of many rollouts as well as other performance metrics.

Finally, we show that even task-specific RL tuning on a single domain pays Sharpening Tax. To mitigate this, we present posterior-tempered group sampling (PTGS), a general sampler that balances exploration and exploitation. It adapts the sampling temperature based on per-task difficulty to smooth policy for hard prompts and sharpen it for easy ones.

Related work. Prior observations that base models have a higher performance upper bound (coverage) than post-trained models come almost entirely from math and coding domains. In agentic domains, however, base models have likely seen far less data in those formats, and they are often assumed to lack the instruction-following and tool-calling skills that agentic tasks require (Zeng et al., 2024; Su et al., 2026). Thus, whether previous coverage analysis transfers to agentic domains remains as a non-trivial gap which motivates our work.

Does RL post-training expand or sharpen the reasoning boundary? Whether RL post-training extends the reasoning boundary of the base model is under active debate. One line of work observes that RL-tuned models lose to their base counterparts in pass@K at large K, suggesting that post-training sharpens the policy around behaviors the base model already possesses (Yue et al., 2025; Zhao et al., 2025; Wu et al., 2025). Another line argues that the boundary can genuinely expand, e.g., through prolonged RL training (Liu et al., 2025a) or grokking (Sun et al., 2026), while others attribute the outcome to properties of the training data and task distribution (Zhang et al., 2025; Shen et al., 2026a; Shao et al., 2026). However, the evidence on both sides comes almost entirely from mathematics and coding (Yue et al., 2025; He et al., 2025), where the skills required for realistic agentic tasks are not assessed, and pass@K can be inflated by lucky final answers (Wen et al., 2026; Dragoi et al., 2025). We bring this debate to agentic environments and give it a quantitative diagnosis, Sharpening Tax, which turns manual inspection of test-time scaling visualizations into a measurable and predictable quantity.

Preserving diversity in post-training. Alongside the investigation of the sharpening hypothesis of post-trained policy, preventing RL post-training from collapsing policy diversity has been a popular topic of research. Common objective-side solutions include exploration bonuses and mitigation of entropy-collapse (Yu et al., 2025; He et al., 2025; Song et al., 2025; Gai et al., 2025) and pass@k-aware learning objectives (Walder and Karkhanis, 2025; Tajwar et al., 2026). Another line is the sampling-side solutions adjusting the rollout distribution through temperature scheduling across training or within rollouts (Yang et al., 2025; Liao et al., 2025; Dang et al., 2026), or estimating power distribution (Karan and Du, 2026).

Method. Metrics. LLM reasoning boundaries are commonly probed through pass@k evaluation (Yue et al., 2025; Zhao et al., 2025). For example, an LLM policy generates k independent rollouts for each task xi, and success is evaluated across these trials. Let ni be the total number of rollouts sampled for task xi ∈D, of which ci ≤ni succeed. Given a budget k ≤ni, we consider three complementary metrics, each computed with the standard unbiased estimator (Chen et al., 2021; Yao et al., 2024) below. i pass@ki. Intuitively, pass@1 captures the sampling efficiency of a policy, pass@k its solution coverage under a finite budget, and passk its success reliability across multiple attempts.

Datasets and environments. Most prior work evaluates pass@K reasoning boundary on math and coding benchmarks, which sometimes check only the final outcome. As a result, an incorrect reasoning trajectory that stumbles onto a lucky final answer still counts as a success, inflating pass@K (Wen et al., 2026). We instead evaluate on three agentic benchmarks that require multi-turn tool calling for final goal achievement: BFCL v4 multi-turn base split (Patil et al., 2025), WebShop (Yao et al., 2022), and ACEBench (Chen et al., 2025). In these environments, intermediate actions and state transitions are checked by design, offering a faithful setup for stress testing of the agentic reasoning boundary. See Appendix A for more details.

Models. Our question is whether post-training extends the reasoning boundary of the base model. To explore this at scale, rather than developing the base and post-trained LLMs from scratch, we leverage open-source checkpoint pairs consisting of a pre-trained base model and its post-trained counterpart (e.g., gemma-4-31B vs. gemma-4-31B-it). Specifically, the evaluation spans 14 backbones from Gemma-4 (Gemma Team, 2026), Ministral-3 (Liu et al., 2026), Qwen2.5 (Yang et al., 2024a), and Qwen3.5 (Qwen Team, 2026) families in HuggingFace collections, ranging from 3B to 35B effective parameters (Tab. 3). The exact training recipes of these models are not fully disclosed, so we use the term loosely for RL post-trained models; since SFT and DPO (Rafailov et al., 2023) also sharpen the policy distribution (Huang et al., 2025), our analysis does not hinge on the exact recipe. Meanwhile, for fair comparison across different model backbones, we turn off thinking mode for all post-trained models that support the enforced explicit reasoning (we also ablate thinking mode in Figure 10). See Appendix A.2 for justification and details.

By harness, we mean a dataset-independent, model-agnostic scaffolding that combines a simple system prompt with relaxed tool-calling and parsing interfaces similar to Yang et al. (2024b) or Wang et al. (2024). As shown in Table 1, harness dramatically improves both pass@1 and pass@32 on benchmarks that base models particularly struggle with, while leaving already-manageable tasks largely unchanged. Meanwhile, we observed that the same harness hurts the post-trained models’ performance (Table 4), consistent with evidence that prompt engineering is not uniformly beneficial for advanced models (Wang et al., 2026a). More generally, harness effectiveness depends on how it interacts with post-training (Kim et al., 2026). Unless specified otherwise, we evaluate base models with the harness and post-trained models with the dataset-default scaffolding; see Appendix A.3 for details and full results.

• Raw area scalability. The cumulative performance recovered by additional compute (parallel rollouts) up to K, defined as the area beneath the pass@K ceiling:

• Sharpening Tax. For a given budget K, the deficit between the base and post-trained policies: TaxX(K) = XBase(K) −XPost(K), X ∈{A, S} (3) A positive TaxX(K) means post-training reduces test-time scalability relative to the base model, charging a “tax” on the coverage gains achievable through compute scaling up to K. Next, we show how broadly Sharpening Tax is charged across models and benchmarks (Fig. 6), and why it is a useful diagnostic (Fig. 7). §6 provides deeper theoretical insights for our tax formulation.

Discussion. Base models in the harness handle agentic tasks, catch up to post-trained ones given test-time compute. Figure 2 presents our first observation, which holds consistently across model backbones and benchmarks. Pre-trained models equipped with a simple harness perform agentic reasoning surprisingly well, catching up to their post-trained counterparts as the rollout budget K grows. This is reminiscent of large language monkeys (Borel, 1913; Brown et al., 2024), but now disciplined under the harness.

We see a clear trade-off between the two policies. Posttrained models achieve higher single-sample accuracy (pass@1), but base models scale much more steeply in solution coverage (pass@K). As the budget grows to 128 rollouts per task, the curves of base models cross over and eventually surpass the post-trained ones across most benchmarks, e.g., over 85% pass@128 on WebShop vs. 56% for RL with gemma-4-31B. The natural next question is how post-training reshapes the policy to produce these outcomes. A popular explanation is the sharpening hypothesis: (RL) post-training amplifies a few high-reward behavior modes that the base model already contains, while suppressing the rest. We dive deeper into this observation.

Is sharpening a curse or a blessing? The scaling curves of base and post-trained models cross over, but when does the crossover happen? Figure 3 plots pass@K curves across four model scales of Gemma-4 (4B, 12B, 26B, and 31B). We observe that larger models pull the crossover point k∗to a smaller rollout budget. In Meanwhile, this bimodalization also makes the policy act more consistently, i.e., for a given prompt, the rollouts within a group show high agreement. Figure 5 confirms this by showing that, averaged over the 12 model-benchmark pairs (largest backbone per family × three benchmarks), the post-trained policy’s coverage (pass@K) and consistency (passK) curves stay comparatively close, whereas the base model shows a remarkably wide gap, i.e., its consistency collapses to zero while its coverage eventually surpasses RL. In short, post-training buys sampling efficiency and consistency by paying with coverage. These phenomena replicate across all model families (See Appendix Figure 18– 21), with Qwen3.5 exhibiting the softest sharpening and Gemma-4 exhibiting the steepest sharpening.

Takeaways. (1) With a simple harness, pre-trained base models are capable agentic reasoners and eventually surpass their post-trained counterparts in solution coverage; (2) the crossover budget at which this happens shrinks with model scale, so whether sharpening helps or hurts must be judged jointly with model scale and test-time budget; (3) the underlying mechanism is distribution sharpening, i.e., post-training bimodalizes the per-task success rate distribution to extrema.

So far, we have examined how RL post-training reshapes test-time scaling dynamics by visualizing the scaling curve. However, manually inspecting full pass@K curves for every model and dataset quickly becomes neither practical nor scalable. Meanwhile, evaluating models with only the endpoint metrics such as pass@K and passK at a specific K says little about the full landscape of the scaling behavior, e.g., whether success saturates within a few attempts or keeps rising. To this end, we propose Sharpening Tax, a diagnostic metric that summarizes the effect of post-training on test-time scalability in a single number. For an evaluation budget K, we define the following quantities.

SharpeningTaxischargedbroadlyacrossmodernpost-trainingpipelines. Figure 6 reports both tax variants’ values of the largest and smallest backbone of each model family on the three benchmarks as a function of the rollout budget k (Figure 26 provides the full results).

Conclusion. We close with a simple lesson on the road to Super Intelligence (The White House, 2026). Base models keep getting stronger (Karan and Du, 2026), and more compute is moving to test time for search (Yao et al., 2023), repeated sampling (Brown et al., 2024), and verification (Kwok et al., 2026). Both trends raise the price of sharpening, since a stronger base model holds more rare but high-value behaviors, and a larger test-time budget is exactly what turns them into solutions. Many challenges we hope future systems will solve, such as open scientific questions (Gottweis et al., 2026) and Millennium Prize Problems (OpenAI, 2026), may be cracked by an unusual shot that works once rather than by a routine that works every time. Post-training should thus be judged not only by its pass@1 but also by how much of the base model’s potential it keeps, e.g., by reporting Sharpening Tax alongside accuracy. In conclusion, reliability and reach should grow together. Whether heavier RL truly moves beyond elicitation remains open, and base-anchored diagnostics like Sharpening Tax offer a direct way to check it as post-training continues to scale.

Limitations. Note that the representative open-source models that we extensively analyze in §3 do not disclose the exact data used in pre-training and post-training, so our analysis remains observational rather than interventional.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Does pretraining establish the ceiling for what reward learning can improve? Can minimal training unlock latent reasoning already present in base models? How do reward signal properties affect model reasoning and safety?