Public articles linked to the same research event.
arXiv Measuring attention-head behavior across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), the authors find most heads change their reading behavior at least once during generation and introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state; at matched memory or budget WakeKV consistently improves miss rate over frozen classification and destructive eviction, and a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality.
Measuring attention-head behavior across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), the authors find most heads change their reading behavior at least once during generation and introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state; at matched memory or budget WakeKV consistently improves miss rate over frozen classification and destructive eviction, and a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality.
Measuring attention-head behavior across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), the authors find most heads change their reading behavior at least once during generation and introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state; at matched memory or budget WakeKV consistently improves miss rate over frozen classification and destructive eviction, and a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality.
Measuring attention-head behavior across three models (1.5B–8B) and three regimes (needle retrieval, long chain-of-thought, multi-turn recall), the authors find most heads change their reading behavior at least once during generation and introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state; at matched memory or budget WakeKV consistently improves miss rate over frozen classification and destructive eviction, and a FlexiCache/vLLM implementation on Mistral-7B improves throughput while retaining LongBench quality.