MILO cuts many-shot context KV cache by up to 50% via block-wise low-rank compression, lifting Qwen2.5 throughput 1.8x with near-flat classification and reasoning performance
Synopsis
The work proposes MILO, a compression framework that exploits low-rank redundancy in many-shot contexts by compressing the KV cache at block granularity and dynamically allocating rank budgets based on information entropy, achieving up to 50% KV cache memory reduction and 1.8x throughput improvement on Qwen2.5 models with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
Figure 1: Top-1% Effective rank of pre-RoPE key and value over different numbers of shots on the MathQA dataset.
arXivInterpretation
MILO introduces a block-wise low-rank compression strategy that compresses the KV cache at block granularity, where each block contains multiple many-shot examples. Prior compression approaches were not designed around block granularity for the many-shot context setting; MILO aligns the compression unit with how many-shot examples are organized. The strategy is presented as a method description, motivated by the low-rank redundancy inherent in many-shot contexts; implementation details are not expanded in the provided abstract text.
MILO dynamically allocates rank budgets based on information entropy to handle heterogeneous context density across blocks, preserving the fidelity of critical blocks while aggressively compressing redundant ones. Compared with a uniform rank budget, tying budget allocation to per-block information entropy links compression strength to each block's information content. The abstract states the design motivation and mechanism but does not provide entropy thresholds or budget ranges.
Experiments on Qwen2.5 models show MILO reduces KV cache memory by up to 50% and improves throughput by 1.8x, with negligible performance degradation on classification and reasoning benchmarks and significant gains over prior baselines. The result reframes the many-shot ICL efficiency bottleneck from KV cache memory into a compressible object and quantifies memory and throughput gains. Evidence comes from experiments on Qwen2.5 models, with the abstract reporting memory, throughput, and benchmark performance; specific benchmark names, example counts, and the baseline list are not given.
Perspective
The work targets long-context inference conditioned on thousands of demonstration examples, especially online serving and on-device deployment where the KV cache is the bottleneck; it applies to LLM inference pipelines using many-shot ICL, validated on Qwen2.5 models. For engineering and deployment readers seeking to lower long-context inference memory and raise throughput, the result offers a reference direction: set the compression unit to blocks containing multiple examples and let rank budgets vary with each block's information entropy.
The provided text is abstract-level material and does not include specific benchmark names, example counts, the baseline list, or ablation results, so the stability of gains across tasks and example scales cannot be judged; how block size is chosen, how information entropy is computed and mapped to rank budgets, and whether compression affects long-range dependencies are open questions requiring the full paper.
