Skip to main content
Back to timeline
arXivSource publication:

WakeKV replaces frozen KV compression with recoverable CPU residency, cutting miss rate at matched memory and raising throughput

Related research and updates

Synopsis

Measuring attention-head behavior across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), the authors find most heads change their reading behavior at least once during generation and introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state; at matched memory or budget WakeKV consistently improves miss rate over frozen classification and destructive eviction, and a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality.

Source-provided article image: WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
Figure 2 ·

Figure 2: Memory-vs-miss-rate Pareto fronts across NIAH, CoT, and multi-turn. Reactive sits at or below frozen’s at all 17 matched-memory points and below evict’s at 19 of 20 matched budgets (Table 1 ).

arXiv

Interpretation

Most attention heads change their reading behavior during generation rather than keeping a fixed role. Most KV-cache compression methods classify heads once, offline or during prefill, and keep that classification fixed throughout generation; this work turns head-behavior change into a measured, reported phenomenon across models and regimes. Head behavior is measured across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), covering four model-regime combinations, with most heads reported to change reading behavior at least once.

Moving cooling heads to a recoverable CPU reservoir lowers miss rate more than frozen classification or permanent eviction. WakeKV makes residency reactive and reversible: state is neither frozen nor destroyed but held recoverably, so a head that becomes relevant again can have its state restored. At matched memory or budget, evaluated across five model-regime combinations, WakeKV consistently improves miss rate over frozen classification and destructive eviction.

Against three cited baselines, WakeKV retains an advantage across comparable combinations. The comparison set is SnapKV, uniform R-KV, and ReasonAlloc, spanning four eligible combinations. The abstract reports comparisons against three baselines across four eligible combinations but gives no specific numbers, effect sizes, or statistical tests.

On real hardware, a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality. The reactive, reversible residency policy is carried into an actual inference stack (FlexiCache/vLLM) rather than remaining a simulation or offline analysis. The FlexiCache/vLLM implementation on Mistral-7B confirms the benefit, reporting improved throughput with LongBench quality retained; the abstract gives no specific throughput figures or quality scores.

Perspective

The work targets long-context generation settings, especially needle retrieval, long chain-of-thought, and multi-turn recall, where head reading behavior changes over time; its audience is inference-systems researchers and engineers concerned with KV-cache memory/budget and throughput. The method is designed to plug into existing inference stacks, with the abstract's concrete instantiation being a FlexiCache/vLLM implementation on Mistral-7B. Its conclusions apply to the three models (1.5B–8B), three regimes, and five model-regime combinations described in the abstract, plus the four combinations comparable with the baselines.

At the abstract level, no specific miss-rate values, throughput gains, or LongBench quality scores are given, and the CPU reservoir's migration overhead, latency impact, and memory footprint are not detailed; the criteria and statistics for judging head-behavior change are likewise not expanded. Readers interested in these quantitative results and overhead trade-offs will need the tables and experimental setup in the full text. In addition, the abstract mentions three baselines but only four comparable combinations, without explaining which combinations are not comparable or why.

Sources