Skip to main content
Back to timeline
arXivSource publication:

CDB lets looped language models infer at token-adaptive depth, realizing about 99% of the early-exit speedup available on Ouro 1.4B

Synopsis

The work introduces and first implements end-to-end continuous depth batching (CDB) for looped language models: it splits inference into prefill, prelude, recurrent-core, and coda stage queues, reforms batches between loop steps, manages a depth-indexed KV cache, and uses a lookahead gate to predict exits one step ahead so tokens with different loop depths can share a forward pass; on Ouro 1.4B and Huginn 3.5B, CDB realizes about 99% and 83%–96% of the estimated maximum speedup, showing fully looped architectures are best suited to depth-adaptive inference.

AI-generated editorial illustration: Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

Interpretation

CDB is the first end-to-end efficient inference implementation for depth-adaptive looped language models, decomposing decode into prefill, prelude, recurrent-core, and coda stage queues that can be batched separately and reforming batches between loop steps. CDB had previously been introduced theoretically with simulated throughput estimates, and later evaluated in ablations that excluded prefill, scheduling, and KV-cache updates; this work includes all these costs and provides the first full implementation and evaluation. On a single 80 GB H100, using one shared codebase, kernels, and admission policy built on Hugging Face Transformers, vLLM, and SGLang, it benchmarks Ouro 1.4B and Huginn 3.5B in offline throughput and online serving, with scheduling traces and latency fits.

The work explains with a FLOP-based bound and a hardware-aware recurrent-step latency model when early exit actually saves time: exits cut recurrent-core FLOPs, but in the memory-bound regime shrinking the batch barely reduces latency, making refill mode important. It corrects decode-FLOP-only speedup estimates into an end-to-end estimate that accounts for the prefill share, and identifies the saturation batch size as the transition between memory-bound and compute-bound regimes. It derives a decode speedup bound and a prefill-corrected end-to-end bound; measured prefill accounts for 3.7%–36% of GPU time, largest on long-context ArXiv, pulling end-to-end speedup well below the decode bound there.

Architecture determines the payoff of depth-adaptive inference: the more expensive the boundary stages (token embedding, LM head, unshared transformer layers), the harder refill becomes, favoring fully looped models. It offers concrete inference-aware guidance for looped architecture design, noting that non-looped layers may improve accuracy yet slow down and complicate scheduling. Ouro 1.4B is fully looped (0-4-0) and refill stays within a small margin of the bound across all three batch-size sweeps, with its largest advantage around or below the saturation batch size; Huginn 3.5B is 2-4-2 with 50% of transformer layers outside the core, where no-refill performs better and the two modes converge as batch size grows.

Asynchronous scheduling plus a lookahead gate reduce GPU idle time to nearly zero without losing accuracy. Exit signals were previously available only after a loop step finished, leaving the GPU idle while the scheduler processed exits; the lookahead gate decides exits one step earlier, giving the CPU a full loop step to prepare the next batch. A gate-distilled lookahead gate closely matches the teacher gate's exit behavior, accuracy, and threshold curves; for Huginn the delayed exit needs only threshold recalibration and incurs no accuracy cost, at the price of constraining the minimum exit depth.

Perspective

The result targets serving systems for looped language models: on a single 80 GB H100, for fully looped Ouro 1.4B and Huginn 3.5B with transformer boundary layers, offline throughput and online serving benchmarks show CDB turns early-exit compute savings into wall-clock speedup. It applies to inference settings that want to allocate loop depth per token and keep batches large in the memory-bound regime; the largest gains go to fully looped models whose boundary stages contain only the embedding and LM head. The method does not change model predictions or exit decisions, so it can work with existing exit gates, paged KV caches, and FlashAttention kernels.

The accuracy cost of KV cache sharing still depends on the model: at fixed depth, Huginn is insensitive to the sharing layout and even does better under single-slot sharing, while Ouro nearly collapses under single-slot sharing, and the authors could not reproduce the near-lossless sharing reported in the original Ouro paper. The lookahead gate constrains the minimum exit depth to 2, which would limit available speedup for models that produce useful outputs after a single loop. The authors evaluate only two looped architectures on a single H100 and exclude multi-GPU execution, preemption, prefix caching, and speculative decoding, which may shift the prefill-decode balance. In addition, this is a full-text read, but the complete online serving results in Appendix D and some figures are not expanded in the text, so cross-workload online serving details remain an open question.

Sources