LEAP trains a 0.6B drafter on target action sequences, making LLM agents up to 60% faster end-to-end with no systematic change in task success
Related research and updatesSynopsis
The authors develop a latency framework for the speculative action round that compares what a round gains (how well the drafter predicts the target and how many steps the task can take before it ends) against what it costs (drafting, waiting for target verification, and executing tools), and use it to design LEAP: keeping the drafter small (0.6B) and making it accurate by training on target action sequences, which yields up to 60% faster end-to-end wall clock time with no systematic change in task success across various datasets, target models, and draft models, while the draft model can also be online trained with no prior trace collection and match offline training performance.
Figure 3: Draft-size ablation, Qwen3 target. Measured speedup, independent planning repeats, and the predictions used in Fig. 4 b. BFCL has one measured size study and no offline-hit prediction. Agreement and cost values are in Table 14 ; pooled results are in Table 1 .
arXivInterpretation
The paper proposes a latency framework for the speculative action round that explains end-to-end speedup through a gain-versus-cost comparison. Prior work largely transplanted speculative decoding ideas and used off-the-shelf models to draft actions, without systematically asking what determines the end-to-end speedup of action speculation. The abstract presents this as a framework analysis and states that across various datasets, target models, and draft models it accounts for most of the measured speedups.
LEAP keeps the drafter small and improves its agreement with the target model by training it on target action sequences. Unlike the existing trade-off where large drafters match the target more often but take longer while small off-the-shelf models are fast but rarely make the same decision, LEAP aims for both low drafting latency and high agreement. The abstract reports that with a small 0.6B model, LEAP agrees with the target on most decisions.
LEAP makes agents up to 60% faster in end-to-end wall clock time with no systematic change in task success. It moves action-phase speculation from a concept toward a measured end-to-end latency benefit while reporting task success as a parallel dimension. The abstract reports up to 60% end-to-end wall clock speedup and states no systematic change in task success; specific datasets, tasks, and statistical details are not given in the abstract.
The draft model can be online trained with no prior trace collection and match the performance of offline training. This lowers the deployment barrier, since target-model historical traces are not required before training the drafter. The abstract states online training matches offline training performance but gives no concrete experimental scale or comparison numbers.
Perspective
The work targets LLM agent rollout settings that require multi-step reasoning and tool calls, and is meant for deployers who want to reduce end-to-end wall clock time without changing the target model's decisions. Its latency framework can be used to judge whether action speculation is worthwhile given a drafter's accuracy and the task's step count; the online training path targets real deployment settings where target traces cannot be collected in advance. The results apply to the datasets, target models, and draft models described in the abstract.
The abstract does not name the specific datasets, task types, target model sizes, or per-decision agreement rates between drafter and target, nor does it state the tasks and hardware conditions behind the 60% speedup, so it is hard to judge under which task lengths and tool-call costs that speedup holds. The criteria and experimental setup for 'online training matches offline training' are likewise not expanded in the abstract. In addition, this reading is limited to the abstract, so figures, ablations, and statistical details in the full text are not included, and the robustness of the reported numbers remains to be checked against the original.
