Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

LEAP trains a 0.6B drafter on target action sequences, making LLM agents up to 60% faster end-to-end with no systematic change in task success

The authors develop a latency framework for the speculative action round that compares what a round gains (how well the drafter predicts the target and how many steps the task can take before it ends) against what it costs (drafting, waiting for target verification, and executing tools), and use it to design LEAP: keeping the drafter small (0.6B) and making it accurate by training on target action sequences, which yields up to 60% faster end-to-end wall clock time with no systematic change in task success across various datasets, target models, and draft models, while the draft model can also be online trained with no prior trace collection and match offline training performance.