Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

iS-KV compresses KV cache online via block-incremental SVD, retaining 82.6% accuracy at 4.06x compression on DeepSeek-R1-Distill-Llama-8B

The work proposes iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning that keeps a recent window exact, incrementally folds older states into bounded-rank representations, and synchronizes historical coordinates as the low-rank basis evolves to maintain representation consistency; on DeepSeek-R1-Distill-Llama-8B it reaches 82.6% accuracy at 4.06-fold persistent-KV compression (original model 83.6%), and on Qwen3-8B it reaches 89.2% accuracy at 5.64-fold compression, consistently outperforming token-eviction baselines under matched memory budgets.