Skip to main content
Back to timeline
arXivSource publication:

WUSH-KV cuts 2-bit KV-cache perplexity from OSCAR's 13.74 to 10.51 and leads OSCAR on all four 8B downstream tasks

Synopsis

The work adapts the WUSH transform, originally for weight-activation quantization, to the KV cache of grouped-query attention: calibration data builds separate key and value transforms per KV head (the value transform folded into the weights, the key transform applied after RoPE), the transform is shown to be near-optimal among sensitivity-balanced transforms for the QuEST quantizer, and end-to-end evaluation in SGLang with OSCAR-style percentile-clipped affine quantization gives 2-bit WikiText-2 perplexity of 10.51 on Qwen3-8B versus OSCAR's 13.74, with the 8B model scoring higher than OSCAR on AIME 2025, MATH-500, GPQA Diamond, and LiveCodeBench v6.

AI-generated editorial illustration: WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Interpretation

It extends WUSH from weight-activation quantization to KV-cache quantization, building a separate key transform and value transform for each KV head: the key transform accounts for the query heads that consume the cached keys, the value transform for the output-projection blocks, both constructed in closed form from calibration-time statistics. Earlier transform-based methods (QuaRot, SpinQuant, FlatQuant, RotateKV, TurboQuant-style Hadamard) mostly use orthogonal or random rotations, and the concurrent OSCAR also restricts its key and value transforms to be orthogonal; WUSH-KV allows general invertible transforms, permitting anisotropic rescaling that jointly balances cache statistics and downstream sensitivity. The paper gives the closed-form construction (Equations 1 and 4) and an offline calibration algorithm, and runs a controlled ablation over all 36 attention modules of Qwen3-8B, calibrated on 128 FineWeb-Edu sequences and evaluated on 32.

For the QuEST quantizer it proves a clipping-aware near-optimality result: under an additive rounding-noise model with normalized-tail and clipping-alignment conditions, the ideal WUSH transform is optimal among sensitivity-balanced invertible transforms up to a constant factor independent of bitwidth. Prior transform work relies largely on empirical comparison without a characterization for clipped quantizers; this work combines the unclipped optimum (Theorem 3) with upper and lower bounds on clipping error and gives an explicit finite-bit bound. Theorem 2 and its proof in Appendix B rest on explicit stated assumptions, and the authors note the guarantee concerns the ideal transform while experiments use the damped version.

Controlled reconstruction experiments show WUSH lowers attention-stage error: at 2-bit with keys and values both quantized, geometric-mean module-output error is at the WUSH and WUSH-A level, below Hadamard and OSCAR; in perplexity experiments WUSH attains the lowest perplexity at every bitwidth, 10.51 versus OSCAR's 13.74 at 2-bit. The experiment holds the quantizer fixed to QuEST, separating transform quality from quantizer design, which is a cleaner control than mixed comparisons in earlier work. Reconstruction covers 36 modules at 2, 3, and 4 bits; perplexity is computed on WikiText-2 with non-overlapping 2048-length sequences processed in 16-token forward chunks, calibrated on 128 FineWeb-Edu sequences of length 32,768, taking about 12 minutes and peaking at 39 GiB on a single NVIDIA L40S.

Integrated end-to-end into SGLang with OSCAR-style percentile-clipped affine quantization, WUSH-KV matches or beats OSCAR at 2-bit: it scores higher on all four 8B benchmarks, is close at 4B and 32B, and is substantially better where OSCAR degrades sharply (LiveCodeBench v6 at 8B and 32B); on long-context RULER NIAH and MRCR it retains higher scores at the longest tested lengths. This is a direct comparison under the same quantizer and cache policy as OSCAR, and the authors state they did not retune clipping ratios for WUSH, making the setup relatively favorable to OSCAR. Evaluation covers Qwen3-4B-Thinking-2507, Qwen3-8B, and Qwen3-32B on AIME 2025, MATH-500, GPQA Diamond, and LiveCodeBench v6 with stochastic sampling over three seeds; the long-context part reports RULER NIAH from 4k to 128k and MRCR from 0k to 128k in segments.

Perspective

The result targets large language model serving with grouped-query attention, offline calibration, and a runtime that supports paged KV caches such as SGLang; it is most directly usable by engineering teams that want to compress the KV cache to 2 bits while preserving reasoning and code-task quality. The method is decoupled from the scalar quantizer and can be paired with per-token clipped quantizers; the value-side transform adds no online cost once folded into the weights, while the key-side transform is applied after RoPE and therefore runs online. The theoretical guarantee concerns the QuEST quantizer and the class of sensitivity-balanced transforms, whereas the end-to-end evaluation uses the OSCAR-style percentile-clipped affine quantizer.

The key-side transform must remain online, which the authors list as extra inference computation and pair with future work on cheaper structured transforms and system-level throughput evaluation; the end-to-end evaluation reuses OSCAR's clipping ratios without retuning them for WUSH, which the authors flag as possible remaining headroom; some OSCAR baseline scores could not be exactly reproduced, so the comparison is rerun under the same SGLang pipeline; MRCR is difficult for all quantized methods and remains well below BF16; and this load is the paper's full text, where figures appear as textual descriptions, so exact curve values should be checked against Figures 1 to 3 in the original.

Sources