TopK-Guided combines bounded token-level sparsity with sensitivity-aware block budgets, improving perplexity and downstream accuracy on Llama-2 and Llama-3
Synopsis
TopK-Guided is a training-free activation sparsity method that combines bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation, consistently improving perplexity and downstream accuracy over TEAL and WINA across Llama-2 and Llama-3 models while preserving essentially the same sparsity-dependent projection compute as WINA, with the largest gains at high sparsity; ablations show the two components provide complementary improvements.
Figure 1: Per-block sensitivity e i e_{i} (Eq. 4 ) on Llama-2-7B and Llama-3-8B, shown at four probe levels s ¯ ∈ { 0.2 , 0.4 , 0.6 , 0.8 } \bar{s}\in\{0.2,0.4,0.6,0.8\} .
arXivInterpretation
Introduces TopK-Guided, a training-free activation sparsity method that bounds token-level sparsity adaptation and allocates sparsity budgets across transformer blocks according to their sensitivity. Existing training-free methods make different trade-offs: threshold-based methods such as TEAL adapt sparsity per token but do not tightly control realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same budget for every token; both also apply the same budget across transformer blocks despite large differences in block sensitivity. The abstract reports consistent improvements over TEAL and WINA on Llama-2 and Llama-3 models and states that ablations show the two components are complementary; specific numbers, model sizes, and evaluation sets are not given in the abstract.
TopK-Guided improves both perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsity-dependent projection compute as WINA, with the largest gains at high sparsity. Decoupling sparsity control from block-level budget allocation yields quality gains without added projection compute, indicating that a fixed per-token budget and a uniform block budget are not necessary trade-offs. Abstract-level comparison covering Llama-2 and Llama-3; the abstract does not report specific perplexity values, accuracy numbers, or sparsity levels.
Ablations show that the token-level bounded adaptation and the block-level sensitivity-aware budget provide complementary improvements. This indicates the limitations of the two prior method families are addressed by distinct components rather than by a single mechanism. The abstract states the ablation conclusion only, without listing ablation configurations, metrics, or values.
Perspective
The method targets training-free LLM inference acceleration, suited to researchers and engineering teams who want to reduce sparsity-dependent projection compute without retraining; the abstract indicates validation on Llama-2 and Llama-3 models and emphasizes improving quality while preserving essentially the same sparsity-dependent projection compute as WINA, so its positioning is quality improvement at an equal compute budget rather than changing the compute itself. High-sparsity settings are where gains are largest, making it relevant for memory- or latency-constrained deployments.
The abstract does not report specific perplexity or downstream accuracy values, the model sizes used, the evaluation datasets, or the sparsity levels, so the magnitude of improvement and the applicable range still require the main tables; how block sensitivity is measured and how budgets are allocated, and whether the method was validated beyond the Llama family or on different hardware, remain open questions.
