Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

SpAx triples weight-read options to speed LLM decoding on offloaded weights by up to 5.57x with at most 10% WikiText-2 perplexity increase

The work introduces SpAx, which replaces the binary choice of whether to read a weight under activation sparsity with three options—fully skipping weights whose activations are closest to zero, reading compressed approximate weights for smaller-magnitude activations, and reading original weights for the largest-magnitude activations—thereby speeding up decoding when weights are offloaded: with weights offloaded to CPU memory, 3.86x average (up to 5.57x) for 16-bit weights and 2.06x average (up to 2.74x) for 4-bit weights; with weights offloaded to flash storage, 3.31x average (up to 4.81x) and 1.54x average (up to 2.03x), at a WikiText-2 perplexity increase of at most 10%.