EDR: training parallel speculative drafters by directly minimizing expected decoding rounds lifts mean accepted length for DSpark and DFly across nine benchmarks
Synopsis
This work formulates speculative decoding with parallel and semi-autoregressive drafters as a Markov reward process and derives Expected Decoding Rounds (EDR), an objective that weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds, together with an exact temporal-difference gradient and an offline evaluator; finetuning DSpark and DFly with EDR consistently improves mean accepted length and outperforms existing objectives across nine math, code, and chat benchmarks.
Interpretation
The paper formulates speculative decoding, conditioned on the target model's output sequence, as a Markov reward process whose states record where the current round started and which position is being verified, with each rejection costing one round, and shows the expected number of decoding rounds equals the sum of diagonal state occupancies. Prior training objectives for parallel and semi-AR drafters rely on block-local surrogates that ignore cross-round coupling; this framework captures that dependence explicitly and characterizes the expected round count exactly. Stated as propositions and theorems with derivations in the appendix, including the forward occupancy recursion and a Bayes-rule derivation of the rejection-acceptance transitions.
It introduces the EDR objective, which exactly equals the expected number of decoding rounds: EOS-aware local rejection costs (the TV distance between draft and target distributions at a position, excluding EOS and normalized by the target's non-EOS mass) are weighted by state occupancies, with no auxiliary hyperparameters. Existing surrogates (e.g., expected accepted length within a round, hand-crafted distributional weights) do not account for EOS termination and weight terms heuristically; EDR uses the actual drafter-dependent occupancies, and the paper proves that maximizing per-round expected accepted length can be strictly suboptimal. Theorem 1 establishes equivalence to the expected round count; Theorem 3 gives a construction showing a block-locally optimal first-order drafter can be a constant factor worse in global MAL.
It derives a temporal-difference form of the EDR gradient that uses only local derivatives, enabling unbiased stochastic optimization from target-model rollouts, and provides exact offline estimators of expected round counts and position-wise acceptance rates so multiple drafters can be compared on shared target rollouts. Differentiating the occupancy forward recursion directly is expensive; the TD form separates the immediate effect on local rejection cost from the downstream effect of changing acceptance probability, weighted by the value gap, making training more efficient. Theorem 2 gives the gradient form; the offline estimators are described as unbiased for round counts and consistent for MAL and acceptance rate, and they do not depend on the drafter.
On Qwen3-4B with DSpark and Qwen3-8B with DFly, finetuning only drafter parameters for one epoch with EDR yields higher or equal MAL than the E2E objective on all nine benchmarks (GSM8K, MATH, AIME25, MBPP, HumanEval, LCB, MT-Bench, Alpaca, Arena-Hard), with larger gains on AIME25, LiveCodeBench, and chat benchmarks. Relative to the original drafters and the end-to-end TV block-local objective, EDR improves consistently without changing architecture or inference procedure, and the gains transfer across two semi-AR drafter architectures and two target-model scales. All methods for the same target model and sampling configuration are evaluated on exactly the same target trajectories, giving paired comparisons that remove variance from independently sampled evaluation outputs; training uses Open-PerfectBlend with at most 512 training anchors per trajectory.
Perspective
The framework targets parallel and semi-AR drafters with a fixed draft block size and applies in settings where training and evaluation are built on target-model rollouts, for example one-epoch finetuning of an existing public drafter checkpoint; the offline evaluator supports paired comparison of multiple drafters on the same target trajectories. The authors note that under high-concurrency serving, verifying long draft blocks can become inefficient, and extending the framework to adaptive block sizes and studying dynamic per-round block-size selection are stated as future directions.
EDR requires computing occupancy weights and value functions for multiple anchor positions, which the authors note is more computationally expensive than block-local objectives, and more efficient ways to use these computations remain an open question. Table 2 shows MAL varies substantially across datasets, and the authors state that the current work does not provide a complete theoretical characterization of this difficulty. In addition, several formulas, table values, and hyperparameters appear as placeholders in the loaded text, so reproducing exact numbers and settings should rely on the original formulas and appendix.
