Skip to main content
Back to timeline
arXivSource publication:

Researchers find high-BOS sink attention heads are redundant components, and pruning them retains downstream tasks more reliably than weight- and activation-based criteria on Gemma-3, Llama-3.1, and Qwen3

Synopsis

The work identifies the BOS sink phenomenon as a key mechanism driving layer-wise redundancy sensitivity in large language models, showing that attention heads with high BOS sink scores, especially in deeper layers, contribute little to predictive performance and effectively serve as dumping grounds for superfluous attention weights, and introduces a simple pruning strategy that removes high-BOS sink heads; experiments on Gemma-3, Llama-3.1, and Qwen3 show this approach identifies redundant transformer components more reliably than weight- and activation-based criteria in terms of downstream task retention, remaining close to dense baselines at low-to-moderate pruning ratios, and that high-scoring sink heads sustain their focus on BOS as context length grows.

Source-provided article image: Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
Figure 1 ·

Figure 1: Attention weights of a representative <BOS> sink head (L26, H0) in Gemma-3-4B. The heatmap reveals a stark attention sink pattern where query tokens disproportionately attend to the <BOS> token during an MMLU task.

arXiv

Interpretation

The paper establishes the BOS sink phenomenon as a key mechanism driving layer-wise redundancy sensitivity in large language models, showing that attention heads with high BOS sink scores are strongly associated with functional redundancy. Large language models were known to contain significant redundancy, but a systematic explanation for why certain components, particularly in higher layers, are more redundant had remained elusive; this work links attention structural properties to layer-wise redundancy. Based on analysis of BOS sink scores of attention heads, combined with experimental observations on Gemma-3, Llama-3.1, and Qwen3.

Attention heads with high BOS sink scores, especially in deeper layers, contribute little to predictive performance and effectively serve as dumping grounds for superfluous attention weights. Reframes redundancy from a vague parameter-count or activation-magnitude issue into a question of specific attention heads being functionally removable. The paper reports that these heads contribute little to predictive performance and observes this pattern across multiple model families.

Introduces a simple pruning strategy that removes high-BOS sink heads, identifying redundant transformer components more reliably than weight- and activation-based criteria in terms of downstream task retention. Compared with magnitude-based methods, this strategy uses attention structural properties as a more direct basis for compression. Experiments on Gemma-3, Llama-3.1, and Qwen3, remaining close to dense baselines at low-to-moderate pruning ratios.

High-scoring sink heads sustain their focus on BOS as context length grows. Indicates that the sink behavior is not an incidental short-context phenomenon but a structural feature maintained with context length. The paper reports observations as context length grows.

Perspective

The results target researchers and engineering practitioners who need to compress large language models, apply to pruning scenarios at the attention-head level, and are validated on Gemma-3, Llama-3.1, and Qwen3. They suggest ranking candidate redundant heads by BOS sink score and preserving downstream task performance close to dense baselines at low-to-moderate pruning ratios. For analysis work on layer-wise redundancy mechanisms, this structural indicator can also serve as an entry point for observing deep-layer attention behavior.

The abstract does not list specific downstream tasks, pruning ratios, evaluation metrics, or numerical results, making it difficult to judge the actual gap relative to dense baselines and the quantitative comparison with other criteria. The stability of high-BOS sink heads across model families, layer counts, and context lengths, as well as how this criterion compares with weight and activation criteria at higher pruning ratios, still needs the full experimental details. In addition, the association between the BOS sink phenomenon and functional redundancy is observational, and its causal direction and mechanistic explanation await further work.

Sources