Skip to main content
Back to timeline
arXivSource publication:

ByteDance team finds chunked KV-cache compression makes long-context retrieval periodically weak at the compression stride, with up to 40 percentage points between phases in DeepSeek-V4

Synopsis

The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.

AI-generated editorial illustration: Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Interpretation

The paper identifies a previously unreported systematic failure mode: in models using chunked KV-cache compression, the same information can be easy to retrieve at one phase and difficult at another, where phase is a token's position relative to compression-window boundaries; the authors call this phase sensitivity, and its period follows the compression stride. Prior work on long-context failure modes focused on where information sits in the context overall (such as Lost in the Middle) or on the query's position within a segment; this work focuses on the source token's phase, a coordinate that recurs every compression stride. In a code-completion example on DeepSeek-V4-Flash-Base, changing only the length of a decorative docstring (24 to 39 tokens) flips the model between the correct continuation "8" and the incorrect "32" with a period of four, matching DeepSeek-V4's stride S = 4; across four filler families and 64 inputs, 60 rankings follow this pattern, whereas DeepSeek-V3.1-Base, which has no chunked compression, ranks 32 above 8 at only 4 of the same 64 inputs.

In a controlled 128K-token needle-in-a-haystack task, retrieval accuracy across phases differs by up to 40.2 percentage points in the DeepSeek-V4 family, and post-training raises accuracy and narrows the gap but keeps the period. The measurement manipulates phase as the only variable: prompt length, query position, and the mean and variance of the target position are matched across residue groups, so differences between groups come from phase rather than position statistics. Each 128K-token prompt holds 16,000 key-value pairs with the target as the 249th pair; DeepSeek-V4-Flash-Base, Flash-0731, Pro-Base, and Pro-0813 each use 256 prompts per residue group with best-to-worst gaps of 40.23, 19.14, 34.77, and 14.84 percentage points, and DeepSeek-V4.1-Flash, with 2,560 prompts per residue group, has a gap of 6.09 percentage points with all four even residue groups above all four odd ones.

Controlled pretraining shows chunked compression itself is sufficient to produce phase sensitivity: all 23 compressed models show weak phases repeating with the stride, while full-attention baselines stay flat. DeepSeek-V4 differs from a standard transformer in many ways and cannot be retrained without compression, so the authors use a Qwen3-0.6B backbone with chunked compression as the only substantial architectural change and sweep window, stride, KV-head count, gating, positional encoding, and local-memory policy. Strides of 4, 6, 8, and 12 give periods of about 4, 6, 8, and 12; full-attention baselines stay within 6.1 percentage points across positions while compressed models reach gaps of up to 78 percentage points under Prefix Padding; with window 8, stride 8, and 8 KV heads, the compressed model's mean accuracy of 59.4% is close to the baseline's 61.1%, but its worst position scores 9.9% against 58.4%.

Causal interventions and theory give mechanistic clues: different attention heads contribute asymmetrically to retrieval at different phases (phase specialization), compression gates' static preferences align with the phases where heads matter, and in an idealized model gradient flow settles each head on a single phase. Prior work established that heads differ in their retrieval roles; this work shows that in compressed models a head's retrieval contribution also depends on phase and links gate preferences to causal contribution. In a one-KV-head model, per-head mean replacement shows some heads matter at only two or three adjacent phases (layer 9 at phases 5 and 6, layer 10 at 3 and 4, layer 14 at 1 to 3); cyclically shifting gate parameters by one slot moves the weak spots by one (R² = 0.98 for the model shown, 0.66 to 1.00 across five models and shifts of up to three slots); in the idealized induction model, theorems show optimal compression is maximally concentrated and gradient flow fixes each head's selection for every input.

Perspective

The result applies to long-context models that use chunked KV-cache compression, in particular the DeepSeek-V4 and DeepSeek-V4.1 families and the Qwen3-0.6B compressed models the authors pretrained from scratch; the actionable suggestion is to report long-context accuracy by phase, for example by shifting the input by a few tokens while keeping the queried information fixed. For a reader, this means that when evaluating or choosing a compressed model, it is worth looking at the worst phase alongside the mean, because average accuracy close to full attention does not guarantee the absence of systematic positional failures. The authors also note that whether explicit coordination across heads and layers can make retrieval quality more uniform across phases without sacrificing average performance or compression efficiency remains open.

The authors state that the mechanistic conclusions are strongest in the controlled settings studied and do not establish a universal cause of phase sensitivity in large, heterogeneous models; gate cycling does not restore the boundary phase, leaving a unified explanation of boundary and within-window asymmetries unresolved. In the DeepSeek-V4 family needle-in-a-haystack evaluation, changing the outer filler shifts the target and distractor records together, so those comparisons do not vary the target's phase independently of the distractors, and the authors use In-sequence Padding in their pretrained models to test this. On the ablation side, mean replacement preserves a reference first moment but does not guarantee an in-distribution activation or a matched norm, and it does not prove that a KV head exclusively stores the target; selecting the largest loss from a sweep is exploratory, and the authors do not test gate-nominated heads on held-out prompts. In addition, the code-completion examples are shifted variants of one example rather than performance on a code-completion benchmark, and the DeepSeek-V3.1-Base comparison also differs in architecture, training, and inference backend. The idealized theory relies on exact embedding geometry, fixed embeddings, and training only gate parameters, and its dynamics proof guarantees persistent selection without guaranteeing complementary coverage of phases across heads.

Sources