iS-KV compresses KV cache online via block-incremental SVD, retaining 82.6% accuracy at 4.06x compression on DeepSeek-R1-Distill-Llama-8B
Related research and updatesSynopsis
The work proposes iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning that keeps a recent window exact, incrementally folds older states into bounded-rank representations, and synchronizes historical coordinates as the low-rank basis evolves to maintain representation consistency; on DeepSeek-R1-Distill-Llama-8B it reaches 82.6% accuracy at 4.06-fold persistent-KV compression (original model 83.6%), and on Qwen3-8B it reaches 89.2% accuracy at 5.64-fold compression, consistently outperforming token-eviction baselines under matched memory budgets.
Figure 2: Overview of iS -KV. The KV cache grows while iS -KV keeps the recent window of w w tokens exact (teal) and folds older tokens into a compact low-rank history (blue), one U Σ B ⊤ U\Sigma B^{\top} per block. Every b b tokens, the pending block F F (orange) is merged in by a Block Incremental SVD that updates basis and coordinates together and solves only a small core, independent of the history length n n . Each position can be reconstructed for query q t q_{t} attends to all t t positions.
arXivInterpretation
It introduces iS-KV, extending SVD-based low-rank compression from a fixed prompt cache to online decoding by keeping a recent window exact and incrementally folding older states into bounded-rank representations to control KV-cache growth. Existing KV-cache compression typically relies on token eviction, whose irreversible deletion can remove historical states that later reasoning may need to revisit; SVD-based low-rank compression retains all positions but was previously non-trivial to extend to online decoding. The abstract describes the design as online low-rank compression and reports accuracy and compression-fold results on DeepSeek-R1-Distill-Llama-8B and Qwen3-8B.
It identifies a history-drift problem in online low-rank compression: if the basis is updated for new tokens while old tokens keep coordinates in the old basis, the stored history drifts substantially. This observation explains why directly moving SVD compression to online decoding is non-trivial and motivates the method design. The abstract states this finding as 'Through our investigation, we find,' without quantifying the drift.
It synchronizes historical coordinates with the updated basis as the low-rank basis evolves, maintaining representation consistency. This coordinate-synchronization mechanism targets the drift problem, keeping historical representations comparable while the basis is incrementally updated online. The abstract presents it as a core mechanism of the method without ablation or error-analysis details.
It achieves high accuracy retention under high compression on long-horizon reasoning and outperforms token-eviction baselines under matched memory budgets. The results indicate that online low-rank compression is a viable alternative to token eviction that avoids irreversible deletion. 82.6% versus the original model's 83.6% at 4.06-fold compression on DeepSeek-R1-Distill-Llama-8B; 89.2% at 5.64-fold compression on Qwen3-8B; and consistently better than token-eviction baselines under matched memory budgets.
Perspective
The work targets autoregressive decoding in long-horizon reasoning, for deployments that need to control KV-cache memory without irreversibly deleting historical states. The method keeps a recent window exact, folds older states into bounded-rank representations, and synchronizes historical coordinates when the basis is updated. The applicability evidence in the abstract comes from DeepSeek-R1-Distill-Llama-8B and Qwen3-8B, together with the token-eviction baselines compared against them.
The abstract does not specify how compression folds are measured, how memory budgets are matched, which evaluation tasks and sample sizes are used, or provide ablations or drift quantification for coordinate synchronization. Readers should still watch the stability of accuracy and compression ratio at longer decoding lengths or other model scales, and how the low-rank basis update frequency and rank bound are chosen. The loaded text is abstract-level, so figures and experimental details are not included, and those details would affect judgments about the method's scope.
