Skip to main content
Back to timeline
arXivSource publication:

WaveFront Decoding fuses drafting and verification into the same recurrent call, speeding up decoding 2.42x on Ouro-2.6B and 3.54x on Huginn-3.5B

Synopsis

The work introduces WaveFront Decoding (WFD), a training-free self-speculative decoding framework for looped language models that exploits two properties of these architectures, namely that intermediate recurrence outputs provide effective draft predictions and that weight sharing lets token states at different positions and recurrence depths be processed in one batched recurrent-block call, organizing these mixed-depth states into a diagonal wavefront so that drafting and verification proceed concurrently within the same recurrent calls while rejected drafts are corrected using full-depth predictions; across six Spec-Bench task categories it achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.

AI-generated editorial illustration: WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Interpretation

WFD turns drafting and verification in looped language models from two separate phases into one continuous pipeline: new positions are drafted at shallow depth while earlier positions are simultaneously advanced toward full-depth verification, both sharing a single batched recurrent-block call. Prior work had already identified the native draft-verifier decomposition of looped models and instantiated it as a two-phase draft-then-verify (DtV) baseline; WFD's increment is to eliminate the phase boundary so that drafting and verification are batched concurrently within the same calls. The paper presents the scheduling algorithm and a wavefront state illustration, and validates the schedule on two architectural families, Ouro-2.6B (full-stack type) and Huginn-3.5B (P/R/C type), indicating the schedule is not tied to one family.

Across the six Spec-Bench task categories, WFD achieves 2.42x speedup over autoregressive decoding on Ouro-2.6B and 3.54x on Huginn-3.5B, outperforming DtV on every task. Relative to DtV this corresponds to roughly a 27% improvement, and the advantage is more pronounced when acceptance is relatively low, for example on Ouro translation where DtV reaches only 1.06x at an acceptance rate of 0.78 while WFD retains 1.92x. End-to-end throughput was measured on a single NVIDIA RTX A6000 with bf16, greedy decoding, user batch size 1, generation limited to 512 output tokens and prompts limited to 1,024 tokens; task-level acceptance rates range from 0.78 to 0.97, with overall rates of 0.92 for Ouro and 0.94 for Huginn.

The paper develops a decoding cost model showing that WFD's speedup comes from converting serial recurrent-block calls into batched calls rather than reducing the recurrence computation required per valid position, and that WFD reaches the asymptotic per-token latency of an infinitely long DtV draft using only a finite wavefront width. The analysis attributes the speedup to amortizing weight reads in memory-bandwidth-bound decoding, and explains that DtV must trade verification amortization against rejection cost when choosing its draft block length, whereas a WFD rejection flushes at most the in-flight positions bounded by the wavefront width. The analysis rests on a roofline-style arithmetic-intensity estimate and a measured latency breakdown; the paper reports that Ouro-2.6B attains only 18% of the roofline throughput bound, with GPU idle time accounting for 65% of per-token latency mainly from kernel-launch overhead, while Huginn-3.5B attains 74%.

Cross-recurrence KV sharing mitigates the KV traffic that WFD incurs at long context lengths because mixed-depth positions access distinct recurrence-specific KV slots, raising WFD's speedup on Huginn-3.5B from 2.54x to 4.81x on GSM8K and from 2.73x to 4.46x on MATH-500. KV sharing is not required by WFD, but it changes the cost of batching positions at different recurrence depths; the paper evaluates the combination separately and notes it is an approximate extension because mixed-depth positions can observe shared KV states at different update stages than sequential autoregressive execution. Huginn has been shown to tolerate inference-time KV sharing; reducing its 32 recurrence-specific KV slots to 4 or 1 shared slots keeps autoregressive accuracy similar, and WFD's acceptance rate rises from 0.88 to as high as 0.98 on GSM8K and from 0.90 to 0.96 on MATH-500; against autoregressive decoding under the identical KV-sharing configuration, WFD stays within at most a few percentage points in accuracy. Ouro was not trained or validated for cross-recurrence KV sharing and naive post-hoc sharing failed to preserve its baseline accuracy, so its original recurrence-wise KV cache organization is retained.

Perspective

The result targets serving scenarios that use looped language models (full-stack type such as Ouro, P/R/C type such as Huginn), especially small-batch, memory-bandwidth-bound decoding. The method is training-free and requires no auxiliary draft model or weight modification, so it can be layered onto existing checkpoints. Cross-recurrence KV sharing extends the applicable setting to long context: the paper reports that shared-KV configurations support contexts up to 64k tokens, whereas the original recurrence-specific caches run out of memory at 32k tokens on Ouro and 16k tokens on Huginn. The paper lists adaptive draft depths, tree-structured drafting, and disaggregating the P/R/C blocks onto dedicated hardware as future directions, indicating the method is designed to compose with such scheduling policies.

Evaluation is limited to a single NVIDIA RTX A6000 and two public checkpoints, so behavior on other looped models and other hardware remains an open question. WFD with cross-recurrence KV sharing is described by the paper itself as an approximate extension, because mixed-depth positions can observe shared KV states at different update stages; under fp32 without KV sharing WFD and autoregressive decoding produce identical token sequences, whereas with shared KV the two decoders no longer match exactly, though accuracy differs by at most a few percentage points. The accuracy differences observed without KV sharing under bf16 are attributed to finite-precision effects from distinct batch shapes and reduction order rather than an algorithmic change in output quality. In addition, on Ouro autoregressive decoding attains only 18% of the roofline throughput bound, with GPU idle time accounting for 65% of per-token latency, so measured speedups on that model are affected by kernel-launch overhead; with CUDA graph capture autoregressive throughput improves by 148%, and WFD's speedup against the optimized autoregressive baseline is 2.30x rather than the 2.78x seen in eager mode. The interaction between adaptive recurrence depth and WFD is also not yet evaluated empirically: reducing average recurrence depth narrows the active wavefront and may reduce WFD's relative speedup over autoregressive decoding, even if absolute latency improves.

Sources