Pre-trained models with a light inference harness cover more solutions on agentic tasks, while post-training imposes a measurable sharpening tax
Related research and updatesSynopsis
The work finds that pre-trained LLMs with a light inference harness, despite far lower single-shot accuracy (pass@1), often surpass their post-trained counterparts in solution coverage (pass@K) given sufficient test-time budget; it proposes the Sharpening Tax metric to quantify post-training's loss in test-time scalability, shows across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases) that the tax is prevalent, estimable from a few rollouts, and correlates with other metrics, and presents PTGS, a difficulty-adaptive Bayesian sampler that reduces the tax.
Figure 37: The difficulty estimate of PTGS tracks the realized success at the end of training. Each point is one of the last 50 training steps of one of five PPO w/ PTGS runs per environment. The x x -axis is the mean posterior estimate p ¯ x \bar{p}_{x} over the 8 prompts of the step, recorded before the rollouts are drawn, and the y y -axis is the realized success rate of its 128 rollouts.
arXivInterpretation
Pre-trained LLMs equipped with a light inference harness can serve as capable agents, and given sufficient test-time budget their solution coverage (pass@K) often exceeds that of post-trained counterparts, despite far lower single-shot accuracy (pass@1). The prior hypothesis that RL post-training merely sharpens base-model behavior had been observed mainly in math and coding tasks; this work extends the test to agentic tasks involving multi-turn tool use and interaction and reports a counterintuitive result. Based on comparisons across 14 base/post-trained model pairs from four families and three agentic benchmarks, 42 cases in total, with the phenomenon reported in most settings.
Post-training pushes tasks toward two extremes, always solved or never solved, thereby improving sampling efficiency and consistency at the cost of solution coverage. The authors offer a mechanism-level account that characterizes post-training as polarization of the task-difficulty distribution rather than a simple gain or loss. The abstract states that the underlying mechanism was analyzed and the metric derived from it; the analysis details are not expanded in the abstract.
The Sharpening Tax is proposed as a diagnostic metric that quantifies the loss in test-time scalability after post-training. The metric turns the solution-coverage cost of post-training into a measurable quantity that can be estimated from a few rollouts and correlates well with other metrics. Across the 42 cases the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics.
Posterior-tempered group sampling (PTGS) is presented as a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty; applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy. It introduces difficulty-adaptive temperature control into the RL training sampling process as a simple pluggable way to reduce the sharpening tax. Compared with a fixed-temperature baseline during RL training in two agentic environments, it reports a smaller tax, higher repeated-sampling solve rates, and higher single-shot accuracy.
Perspective
The results target researchers and engineering teams who post-train LLMs with reinforcement learning and deploy them as agents, apply to agentic tasks with multi-turn tool use and interaction, and require evaluating solution coverage under a sufficient test-time budget. The Sharpening Tax serves as a diagnostic that can be estimated from a few rollouts, suitable for quickly assessing the test-time scalability cost of post-training before training or deployment; PTGS, as a plug-and-play sampler, is meant to replace fixed-temperature sampling during RL training.
The abstract does not give the specific tasks in each benchmark, model scales, concrete pass@1 and pass@K values, the number of rollouts needed to estimate the sharpening tax, or the training details and hyperparameters of PTGS in the two agentic environments; how task polarization is measured in the mechanism analysis is also not expanded. Readers who want to adjust training or inference budgets accordingly still need to consult the experimental setup and ablation results in the full text.
