TIDE reuses inference-time hidden states to adapt draft models online, reaching up to 1.66x throughput and recovering performance when static drafts degrade it
Synopsis
TIDE is a serving-engine-native framework that reuses the target model's intermediate hidden states produced during inference as training signals for online draft adaptation, avoiding additional target model computation and serving-time overhead, and combines this with adaptive runtime control that activates speculation and draft training only when beneficial and with mapping of inference and training onto appropriate GPU classes; across diverse real-world workloads TIDE achieves up to 1.66x throughput over no-speculation baselines, recovers performance on misaligned workloads where static draft models degrade throughput, reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches, and improves system throughput by up to 1.
Fig. 1 : Normalized throughput of speculative decoding with a static draft model versus TIDE across four non-English Alpaca datasets. The static draft model, trained on English data, degrades throughput below the no-speculation baseline on all workloads. TIDE recovers and exceeds baseline throughput by continuously adapting the draft model to the current workload.
arXivInterpretation
TIDE proposes reusing the target model's intermediate hidden states, already produced during inference, as training signals for draft adaptation, enabling online draft adaptation without extra target model computation. Existing draft training typically requires additional target model computation or serving-time overhead; TIDE changes the source of training signal to hidden states that already exist during inference and embeds adaptation in the serving engine itself. The text describes the mechanism as reusing "target model's intermediate hidden states generated during inference as training signals" and states it avoids "additional target model computation and serving-time overhead"; the evidence comes from the framework design plus overall results on real workloads, with no separate ablation for this mechanism reported in the loaded text.
TIDE introduces adaptive runtime control that activates speculation and draft model training only when they are beneficial. A static draft model can slow throughput when workloads evolve, whereas TIDE makes the decision to enable speculation and training a runtime decision rather than a fixed configuration. The text states TIDE "employs adaptive runtime control to activate speculation and draft model training only when beneficial" and supports this with "recovering performance on misaligned workloads where static draft models degrade throughput".
TIDE exploits heterogeneous clusters by mapping inference and training to appropriate GPU classes, improving system throughput by up to 1.22x on heterogeneous GPU clusters. Placing draft-adaptation training and inference on suitable GPU classes turns heterogeneous hardware into usable capacity rather than a burden. The text reports "improves system throughput by up to 1.22$\times$ on heterogeneous GPU clusters", a system-level measured result.
Across diverse real-world workloads, TIDE achieves up to 1.66x throughput over no-speculation baselines and reduces training time by up to 3.02x and storage requirements by 24x compared to existing draft training approaches. These figures span inference throughput, training cost, and storage footprint together, indicating that online adaptation lowers the cost of adaptation itself alongside its benefits. The text reports "up to 1.66$\times$ throughput over no-speculation baselines" and "reduces training time by up to 3.02$\times$ and storage requirements by 24$\times$ compared to existing draft training approaches", all as multiples relative to baselines.
Perspective
This work targets high-performance LLM inference serving systems, especially deployments where workloads change over time and the cluster contains heterogeneous GPU classes. It enables a serving engine to keep adapting the draft model at runtime and to turn off speculation and training when speculation is no longer beneficial; for systems engineers, the directly reusable ideas are reusing inference-time hidden states as training signals and adaptive runtime control. The reported results come from measurements on diverse real-world workloads and heterogeneous GPU clusters, so the intended scope should be read as these serving scenarios rather than any arbitrary inference configuration.
Only abstract-level text was read here, without experimental setup, workload composition, baseline configuration, or ablation results, so it is not possible to attribute each gain to hidden-state reuse, adaptive control, or heterogeneous mapping separately. The reported multiples are all "up to" upper bounds, and typical-workload behavior still needs confirmation in the full text. In addition, how quickly draft adaptation converges when workloads switch rapidly, and how much hidden-state reuse interferes with the target model's own inference, are questions worth watching.
