Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Theoretical analysis of sparse attention: JS divergence is fixed by truncation mass alone, information loss grows with sequence length, and generalization bounds tighten with sparsity

This work presents a unified theoretical analysis of sparse attention: it proves that the Jensen–Shannon divergence between full and sparse attention is determined solely by the discarded truncation mass with a closed-form expression, uses order statistics under i.i.d. sub-Gaussian attention-score assumptions to show that the high-probability lower bound on truncation mass increases with sequence length, and derives sparsity-tightened generalization bounds via Rademacher complexity and a Xu–Raginsky mutual-information bound with entropy and covering-number analysis, verifying the closed-form divergence and the growth of information loss on GPT-2 (124M) with WikiText-103.