Post-training pushes tasks to always-solved or never-solved: a "sharpening tax" appears across 14 base/post-trained pairs on three agentic benchmarks, and PTGS lowers it
Synopsis
The work compares 14 open base/post-trained checkpoint pairs (Gemma-4, Ministral-3, Qwen2.5, Qwen3.5, 3B to 35B) on three multi-turn tool-calling benchmarks (BFCL v4 multi-turn, WebShop, ACEBench) and finds that harness-equipped base models, despite far lower pass@1, often surpass their post-trained counterparts in pass@K given enough test-time budget; it names this coverage loss the Sharpening Tax, shows post-training bimodalizes per-task success toward always-solved or never-solved, and proposes PTGS, a per-prompt difficulty-adaptive temperature sampler that improves single-shot accuracy and coverage while paying a smaller tax in two agentic environments.
Interpretation
On agentic tasks, pre-trained base models with a light harness can perform multi-turn tool use and, as the test-time budget grows, match or surpass their post-trained counterparts in solution coverage (pass@K). Prior evidence for the "post-training merely sharpens, trading coverage for accuracy" hypothesis came almost entirely from math and coding, while agentic tasks with multi-turn tool use and environment interaction were expected to depend on capabilities newly acquired in post-training; this work moves the comparison to three process-scored agentic benchmarks. 14 base/post-trained checkpoint pairs from four families across three benchmarks, 42 model-benchmark combinations, 128 rollouts per task; e.g., on WebShop the gemma-4-31B base model exceeds 85% coverage at large budgets versus 56% for its post-trained counterpart.
Post-training bimodalizes each task's success probability: the middle category ("solvable given compute") shrinks sharply as tasks are pushed to the extremes of always-solved and never-solved, buying sampling efficiency and consistency at the cost of coverage. It turns an observation previously made mainly by visually inspecting scaling curves into a characterization of the per-task success-probability distribution, with an always-pass / pass-given-compute / always-fail breakdown. For gemma-4-31B on WebShop, the middle category drops from 87.6% to 30.0%, always-pass rises from 0.0% to 26.0%, and always-fail rises from 12.4% to 44.0%; the pattern replicates across all four families, with Qwen3.5 sharpening least and Gemma-4 most steeply.
It proposes Sharpening Tax, a scalar diagnostic quantifying the loss in test-time scalability after post-training, and shows the tax is prevalent in most settings, estimable from a few rollouts, and strongly correlated with other metrics. It compresses the pass@K scaling curve, which otherwise requires manual inspection per model and dataset, into a single comparable and predictable number with a probabilistic interpretation in terms of how much success depends on retries. The tax rises with budget in 36 of 42 combinations; a tax estimated from 8 rollouts on half the tasks predicts the tax on the held-out half with high Spearman correlation, and it shows a clear positive rank correlation with the consistency gap.
It proposes posterior-tempered group sampling (PTGS), which estimates each prompt's difficulty online via a Beta-Binomial posterior and adapts the sampling temperature, heating hard prompts and cooling mastered ones; with PPO and GRPO on Sokoban and FrozenLake it improves both pass@1 and pass@K while paying a smaller tax. Unlike global temperature scaling or fixed truncation, PTGS changes only the sampling distribution of training rollouts without touching the RL update, making it plug-and-play, and it comes with a theoretical analysis of the rollout groups RL learns from. Fine-tuning Qwen2.5-7B-Instruct for 200 steps, averaged over five runs: on Sokoban PPO gives pass@1/pass@K of 46.5/55.0 with tax 0.094, while PPO with PTGS gives 61.1/69.7 with tax 0.081; on FrozenLake PPO gives 63.7/74.1 with tax 0.039, while PPO with PTGS gives 65.0/80.0 with tax 0.020.
Perspective
The work targets researchers and engineering teams working with open base/post-trained checkpoint pairs, in settings of text-based LLM agents, multi-turn tool calling, and binary success scoring: BFCL v4 multi-turn, WebShop, and ACEBench, with budget counted in rollouts (128 per task). Its directly usable outputs are a tax metric estimable from a few rollouts, which can route test-time budget between base and post-trained models per task, and PTGS, a sampler that changes only the training rollout temperature and can be plugged into modern RL pipelines such as PPO and GRPO without modifying the RL algorithm. The authors note that because the open models do not disclose their pre-training and post-training data, the analysis is observational rather than interventional, and replicating it in a controlled setup with fully known data mixes would be needed for a causal account of when and how sharpening emerges.
Open questions the authors list include: how different types of post-training (RL, supervised fine-tuning, on-policy distillation) charge the tax differently has not been compared; the scope is text-based LLM agents, so whether the accuracy-consistency-coverage trilemma holds for multimodal LLMs and vision-language-action models remains open; Sharpening Tax relies on a binary success signal, so extending it to open-ended generation would require replacing pass@K with quality-gated diversity measures; and how sharpening reshapes other uses of base/post-trained pairs, such as contrastive decoding, model merging, and process reward modeling, is left to future work. In addition, the loaded text is the full paper but some figures and tables appear as images, so specific numerical details should be checked against the original figures.
