SPIN predicts KV-block importance from past indexer scores, skipping 30–40% of indexer input blocks on DeepSeek-V4 and lifting vLLM throughput by 14.9%
Synopsis
SPIN (Shadow Predictive Indexer) uses per-layer statistics of indexer scores from previous decoding iterations, maintaining EMAs along vertical and diagonal patterns to predict KV-block importance and skip context blocks before the indexer; on DeepSeek-V4 Flash and Pro across long-context and agentic benchmarks it preserves task quality at 30–40% block sparsity, and integrated into vLLM it improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.
Figure 1 : Indexer top- K K patterns from a single DSA layer on a prompt from summscreenfd . The y-axis represents decode iterations. The prompt is 9.7K tokens long, and 1,400 context KV slots are visualized along the x-axis. DeepSeek-V3.2 exhibits prominent diagonal bands and vertical lines, whereas DeepSeek-V4 exhibits mostly vertical patterns.
arXivInterpretation
SPIN targets the indexer overhead itself: rather than changing core attention's fixed top-k budget, it prunes the indexer's input by predicted block importance so that each decoding step no longer scores the full KV cache. Prior indexer acceleration relied largely on cross-layer reuse (e.g., IndexCache, GLM-5.2's inter-layer sharing) or hierarchical filtering (e.g., HISA); SPIN instead explicitly models the temporal patterns of indexer scores, using per-slot and relative-offset EMA states to select input blocks. On DeepSeek-V3.2 and DeepSeek-V4 Flash, five predictors are compared on a 9.7K-token summscreenfd prompt and a 94K-token AA-LCR prompt; prev-iter reaches Pearson correlations of 0.913/0.902 (V3.2) and 0.876/0.849 (V4 Flash), with top-k recall of 0.824/0.615 and 0.751/0.577, above prev-layer and the random baseline.
SPIN preserves task quality at 30–40% block sparsity on long-context and agentic tasks, and on MRCRv2 random exploration narrows the gap to all-keep without increasing the block budget. The work applies sparsity at the KV-block level to fit paged-attention backends, and addresses staleness—skipped blocks receive no score updates—by reserving part of the same block budget for random exploration. On LongBench-v2, scores at 20%–50% target sparsity differ from all-keep by at most 2.6% relative; on AA-LCR by at most 2.4 points (3.7% relative); on RULER at 50% sparsity p99 miss reaches 48.83% while the score drops only 1.24 points; tau2-airline shows no degradation at 30% sparsity and a 4.1% drop at 50%; on MRCRv2 at 256K/512K/1M the gap grows with sparsity without exploration and 5% random exploration substantially narrows it.
SPIN is implemented in vLLM and treats speculative decoding as a first-class consideration: it predicts block masks for multiple verifier queries and unions them, deferring EMA updates until acceptance outcomes are known and replaying prefix-valid rows. This lets indexer sparsification coexist with DeepSeek-V4's native multi-token prediction (MTP), whereas prior work largely focused on the single-query decoding path or on accelerating the post-indexer top-k operator. Offline replay of 60 MTP traces shows prefix-valid update reduces p99 miss by 5.9–11.0% versus row-0 update; online on SPEED-Bench, the chosen strategy over 1,499 matched requests yields average acceptance length 2.4097 and draft-token acceptance rate 47.28%, within 0.0049 and 0.5% of all-keep's 2.4049/47.19%, while retaining 39.51% realized block sparsity.
End-to-end serving gains grow with target sparsity because SPIN's own overhead is largely fixed. This gives deployment a tunable knob for long-cached agentic workloads: once the fixed overhead is paid, further reductions in indexer work translate more directly into throughput. On eight NVIDIA B300 GPUs (TP8/EP8) with 507K cached context tokens, 5K new input tokens, and 96 concurrent requests each generating 16K output tokens, 40% sparsity gives +10.6% output throughput and 9.2% lower median ITL; 50% sparsity gives +14.9% and 13.2%, with the throughput gain rising 40.7% from 40% to 50%.
Perspective
The result targets long-context and agentic inference serving with indexer-based sparse attention (such as the DeepSeek-V4 model class), especially vLLM deployments with long cached prefixes and high-concurrency decoding; the method is training-free, adds no auxiliary layers or networks, applies sparsity at the block level for paged-attention backends, and coexists with DeepSeek-V4's native MTP. For inference engineers and systems researchers seeking to cut very-long-context decode latency without retraining, SPIN offers a reusable design for selecting indexer input blocks by predicted importance within a fixed block budget, along with ablation guidance on block pooling and exploration policies.
SPIN relies on historical indexer scores and therefore has limited ability to anticipate sudden attention shifts not represented in its history; skipped blocks receive no score updates, so predictions and observations can feed back on each other, and random exploration only offers another chance to be observed rather than guaranteeing timely discovery or full task-score recovery. Evaluation centers on DeepSeek-V4 and its native MTP, leaving other indexer-based sparse-attention models and other speculative-decoding methods untested. In addition, three examples in the 1M MRCRv2 bin that exceed the model context limit retain zero scores, and RULER runs each configuration once given its large evaluation set; the specific effect of these settings on score variation is worth noting when reproducing.
