LSD cuts the length-scaling tax on already-solved queries from 19.0% to -3.7% while matching or beating RL Pass@1
Synopsis
The work defines the length-scaling tax (LST) as the excess response length that RL post-training adds on already-solved queries without a matching accuracy gain, and proposes Length Self-Distillation (LSD), which routes solved prompt groups to on-policy distillation while keeping the original RLVR objective for unsolved groups, using an exponential moving average of the same policy lineage as teacher with no external model; LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks while matching or improving average Pass@1 over RL.
Interpretation
The paper introduces and quantifies the length-scaling tax (LST): on a fixed easy-query set, response length on already-solved queries grows during RL post-training without a commensurate accuracy gain. Prior characterizations of reasoning efficiency relied mostly on coarse aggregate statistics such as average response length, and lacked a systematic metric for the unintended lengthening that later RL updates impose on easy queries; LST freezes an easy set at an anchor checkpoint and uses the shortest mean length among checkpoints meeting an accuracy criterion as reference, isolating this training-induced inefficiency. A standard group-relative RLVR baseline initialized from Qwen3-4B-Base, post-trained on deduplicated DAPO-Math-17K and evaluated on AMC 2023, AIME 2025, and AIME 2026; regardless of which anchor defines the easy set, accuracy stays largely stable while response length keeps increasing.
The paper proposes Length Self-Distillation (LSD), which routes by the current rollout group's empirical solve rate: solved groups go to on-policy distillation while unsolved groups retain the original RLVR objective. Existing on-policy distillation is used mainly as a general compression objective, while dynamic sampling discards all-correct groups as uninformative and difficulty-aware curricula downweight easy problems; LSD instead treats all-correct groups, whose group-relative advantage collapses to zero, as the place where a token-level preservation signal is needed. Three implementations are developed: supervised-gradient forward KL (SG-FKL), supervised-gradient reverse KL (SG-RKL), and policy-gradient reverse KL (PG-RKL), analyzed for their different levels of policy-control granularity; the teacher is an exponential moving average checkpoint of the same policy lineage, requiring neither a stronger external teacher nor a separately prompted concise model.
On single-turn mathematical reasoning and multi-turn agentic tasks, LSD substantially curbs easy-query length growth while maintaining or improving benchmark performance. Compared with RL, CRISP, Fixed SG-FKL, and RL + Length Penalty, LSD's gain comes not from shortening responses globally but from concentrating the shortening on already-solved queries. On single-turn reasoning LST falls from 19.0% under RL to -3.7% (SG-FKL), -10.92% (SG-RKL), and 1.45% (PG-RKL), with average Pass@1 of 44.20%, 42.80%, and 43.56% versus 42.90% for RL; on multi-turn agentic tasks LST falls from 31.4% to 13.7%, 9.2%, and 16.1%, where SG-FKL and SG-RKL also reduce average turns and raise Pass@1, while RL + Length Penalty is shortest but drops Pass@1 from 29.50% to 22.58%.
Ablations show that teacher lag and routing threshold jointly set the trade-off between preserving conciseness and acquiring new capabilities. The paper links LSD's two new hyperparameters (EMA half-life and routing threshold) to the token share actually sent to distillation, teacher parameter lag, and distillation loss, showing that stronger preservation does not require greater distillation exposure. A half-life of 4 gives the highest average Pass@1 (41.85%); lowering the routing threshold from 1.0 to 0.85 and 0.75 routes more partially solved groups to distillation and lowers average Pass@1 to 40.91% and 39.31%, while the distillation token share rises from 10.31% to 19.00%.
Perspective
The result targets settings that use verifiable-reward RL post-training and whose training distribution contains both solved and unsolved prompts, such as single-turn mathematical reasoning and multi-turn retrieval agents; it is most directly useful to training teams that want to control reasoning cost without introducing an external teacher model. LST is defined as excess length relative to the shortest mean length meeting an accuracy criterion, so it presumes a freezable easy-query set and an accuracy-qualified reference checkpoint; negative LST means responses are shorter than that reference, and accuracy must be checked separately on the same fixed easy set. The authors list evaluation on larger models as future work.
The easy set is frozen at an anchor checkpoint, so different anchors yield different LST values and cross-paper comparisons should note the anchor and threshold used. The three distillation objectives do not rank consistently: SG-RKL has the lowest LST but also the lowest average Pass@1, and PG-RKL has the highest Pass@1 on multi-turn tasks while using more turns than RL, so lower LST does not automatically mean fewer interactions. Individual cases in the appendix mark the edges of the aggregate result: on aime25_6, SG-RKL and PG-RKL fall below RL in correctness, and Fixed SG-FKL shortens responses while losing substantial accuracy. In the ablations, routing threshold, distillation token share, and effective distillation coefficient move together, and the paper notes these coupled changes do not isolate which training signal causes the performance difference. The training logs also do not provide per-response lengths by route, so a per-response length histogram cannot be reconstructed from the available aggregates.
