MASA recasts sparse-attention selection as matrix approximation, lifting long-context accuracy across three sparse frameworks at no extra budget
Synopsis
The authors argue that existing sparse-attention selectors keep large or high-mass entries of the attention matrix even though Transformers use that matrix through multiplication with value vectors, and propose MASA, which replaces attention-mass ranking with a closed-form score measuring each sparse unit's reduction in matrix-product approximation error; plugged into MInference, SeerAttention, and FlexPrefill without changing their sparse kernels, candidate units, or budgets, it yields consistent accuracy gains on RULER, InfiniteBench, and LongBench with negligible latency overhead.
Figure 1: A minimal two-column counterexample. Under the same one-column budget, conventional mass-based ranking selects u 1 u_{1} , whose mass is 66 × 66\times larger but incurs 98 % 98\% relative matrix-product error. MASA selects u 2 u_{2} and reduces the error to 3 % 3\% . Here, 𝑿 {\bm{X}} is an illustrative probe matrix.
arXivInterpretation
The paper reformulates sparse-attention selection as matrix approximation: under a fixed budget, the retained sparse units should be those that best preserve the matrix product A·V, not those with the largest entries in the attention matrix. Prior methods (MInference, SeerAttention, FlexPrefill and others) ultimately rank by magnitude, keeping large scalar entries or high-mass regions; the paper shows entrywise approximation matches output error only when the attention matrix behaves like an identity over token positions, a condition that does not hold in long-context attention. The paper derives that expanding the matrix-product error leaves entrywise magnitude ranking as only the identity-Gram special case, and gives a minimal two-column instance in which a lower-mass candidate receives a higher MASA score because it preserves the matrix product substantially better.
MASA provides a plug-in closed-form score: each candidate sparse unit is ranked by the reduction in squared matrix-product error obtained by retaining it, reducing to a sum of pair scores for vertical and slash units and retaining an intra-block interaction term for block units. The method keeps the baseline's sparse unit definitions, sparse kernels, and budget rules, replacing only the ranking score, which makes it orthogonal to low-level attention kernels and KV-cache design; scoring happens at the baseline's existing profile resolution, so no new full-resolution dense attention map is required. The paper gives the closed-form pair score, showing the score depends on the probe vector's direction and norm and its alignment with the dense matrix action, and provides a memory-safe matrix-product implementation plus soft norm balancing and norm clipping probe constructions.
Across three sparse-attention baselines, three long-context benchmarks, and both Llama- and Qwen-family backbones, MASA improves all eight matched average scores, with gains most evident at longer context lengths. All comparisons hold the model, sparse budget, sparse kernel, sparse unit family, and data processing fixed and change only the ranking score, isolating the effect of the selection objective; on RULER, for example, MASA moves SeerAttention from below the dense reference to above it, and FlexPrefill-Llama's gain exceeds its original advantage over full attention. MASA improves 11 of 12 matched length-level scores across the four RULER configurations at 32K, 64K, and 128K, with mean gains at 32K–128K exceeding those at 4K–16K for every configuration; Appendix E directly measures lower normalized squared output error in all four framework–backbone pairs, with all four reported paired intervals above zero.
The extra computation from replacing the ranking score is empirically negligible: MASA evaluates the reduction score inside the existing ranking stage rather than adding a new attention component. This means the accuracy gains do not come at the cost of the efficiency advantage of sparse attention, and the overhead stays negligible at 128K, the regime where sparse attention is most needed. The paper reports kernel-level latency normalized to dense FlashAttention-2 at sequence lengths 32K, 64K, and 128K across four framework–backbone combinations.
Perspective
MASA's scope is correcting the selection score within a given sparse-attention framework; it does not search for new sparse kernels, new sparse unit families, or new budget schedules, which the paper presents as what keeps the comparison clean. Its implementation and experiments instantiate MASA in the sparse prefill setting, where the main cost comes from computing attention over a long prompt before decoding, with MInference, SeerAttention, and FlexPrefill as baselines. The intended users are researchers and engineers who want long-context accuracy gains without changing existing sparse kernels or budgets; probe construction (raw values, soft norm balancing, norm clipping) remains an implementation choice, and the paper selects the tested configuration with the highest aggregate score per benchmark.
Readers should still watch that the best probe choice (raw values, soft norm balancing, norm clipping, and its threshold) varies across frameworks and benchmarks, and the paper selects the tested configuration with the highest aggregate score per benchmark, so probe tuning in deployment deserves attention. The Appendix E error measurement uses 64K RULER inputs, 208 prompts per model, and eight prespecified layers, and the paper states the reported intervals describe these selected-setting comparisons rather than an independent held-out evaluation. In addition, several numeric values in the main-text tables are absent from the loaded text, so specific scores can only be checked against the full tables in Appendices D and F; exact reproduction of a given configuration's numbers should rely on those full tables.
