Skip to main content
Back to timeline
arXivSource publication:

Keeping only the top-attention tokens without retraining, nine checkpoints need effective attention sets of 7 to 283 tokens, and longer context raises the required set size

Synopsis

Without retraining, this work retains only the tokens with the highest attention weights at each head, layer, and query while keeping their original weights unchanged, and estimates the effective attention set size needed to stay within a chosen loss tolerance by measuring the increase in negative log-likelihood (NLL); it finds that relatively small selected sets keep NLL close to the full-attention baseline, that attention-based selection substantially outperforms random selection, that the required set size grows with context while its fraction of context decreases, that in BABILong experiments with a fixed annotated supporting fact additional background pushes support tokens down the attention ranking and reduces their attention mass, and that renormalizing the retained weights can subs

AI-generated editorial illustration: Retrieval Capacity of Self-Attention Under Competition

Interpretation

The paper proposes a useful-token framework that distinguishes an unknown useful set from observable selected sets and derives recovery bounds connecting ranking errors to the number of tokens that must be retained. Prior work by Mudarisov et al. observed that sets selected by the largest attention weights show stronger geometric separation than random sets of the same size in the space of weighted value vectors; this work turns that observation into a testable functional framework with a quantitative relation between ranking inversions and the size needed to recover the useful set. Formal propositions with proofs (recovery bounds in Appendix A.1) alongside empirical measurements on nine decoder-only checkpoints; the useful set itself is unobserved, and the experiments examine selected sets and their consequences.

Across nine checkpoints, sets selected by attention or contribution magnitude are geometrically more separated than random sets of the same size, and retaining them produces far less NLL degradation than random selection; the set size needed to keep average loss within a given tolerance nevertheless varies substantially across models. The paper compares geometric separation with loss preservation and shows that geometric separation alone does not establish that model loss is preserved, moving structured selection from a descriptive observation to a functional measurement. Nine checkpoints (Qwen-2.5-1.5B/7B, Gemma-7B, Gemma-2-9B, Llama-2-7B, Llama-3-8B, Llama-3.2-1B, Mistral-7B-v0.3, Mistral-Small-24B-Base-2501), 50 documents per model per corpus on OpenWebText and WikiText-103; Table 2 gives corpus-specific estimates, for example 283 for Gemma-7B at 1% tolerance on OpenWebText, 51 for Mistral-7B, and 12 for Gemma-2-9B at 5% tolerance.

Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range; in BABILong qa1, holding the annotated supporting fact fixed while adding background text moves support tokens down the attention ranking, reduces their attention mass, lowers their recall among the 64 highest-weight tokens, and increases the required set size in several models. Natural context extension changes both useful information and competition, so the paper adds a controlled experiment with fixed annotated support to isolate competition, and reports that individual pairwise outranking probability falls while the growing number of competitors makes overall ranking worse. The matched-target context-length evaluation covers 256 to 2048 with complete results for three models on both corpora (27 of 32 planned conditions complete); BABILong uses 100 QA examples per background with 86 validated, and across models mean support displacement rises by about 3.6x, attention mass falls to about 0.4 of its initial value, and recall@64 drops by 63.8 percentage points.

Renormalizing the retained attention weights can substantially reduce the required set size, showing that effective set size also depends on how selected representations are combined; conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. This extends the explanation from ranking to aggregation and gives an idealized task-loss example in which growing deletion set sizes are needed even when every value vector is identical, purely from preserving output amplitude. Language-model and QA normalization controls (Tables 9 and 10), for example Gemma-7B dropping from 226 to 7 at 5% under attention ranking; the theoretical results are conditional propositions with stated assumptions that the authors say have not been verified for the tested models.

Perspective

This work is aimed at readers studying the attention behavior of long-context language models and at researchers who need to evaluate attention sparsification or cache-reduction schemes. It offers a reusable measurement procedure: estimate the effective attention set size under a given loss tolerance and examine separately the effects of context length, competition, and aggregation. The paper states that the measurement applies to the evaluated models, corpora, and tolerances, and that the intervention computes dense attention before selection, so the results are meant to characterize prediction degradation rather than to deliver an inference speedup.

The useful set is unobserved and attention weights alone do not establish relevance, so the relation between effective set size and the size of a local useful set still depends on ranking errors, the values being combined, and the sensitivity of subsequent computation. The correlation between geometric separation and loss degradation weakens or changes sign once small selected sets are excluded, so geometry should not be read alone as a predictor of functional sufficiency. Annotated BABILong support is only a partial reference for relevance, average loss can conceal changes on individual examples, and preserving a weak full-model baseline does not establish successful task solving. Several conditions remain unresolved or incomplete (for example Gemma at 2K and Mistral at 4K search limits, and missing Mistral-7B WikiText-103 results), and threshold estimates are conditional on the evaluated set sizes. The theoretical models are conditional, their assumptions have not been verified for the tested models, and the paper does not claim a universal scaling law.

Sources