Skip to main content
Back to timeline
arXivSource publication:

PrismQuant rotates dominant activation energy into the null space of grouped quantization, reaching 3.85 perplexity at W4A4KV4 on Llama-3.1-70B, only 0.22 points below full precision

Synopsis

The work introduces PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the group-constant subspace of asymmetric grouped INT4, formulates rotation design as a Ky Fan trace maximization with a provably optimal closed-form solution, realizes it with compact Householder (compact-WY) transforms without gradient training, and evaluates W4A4KV4 post-training quantization on Llama, Qwen, and Mistral (including dense models up to 70B and a 30B mixture-of-experts model): on Llama-3.2-3B it sets the state of the art among compared methods in perplexity and accuracy, on Llama-3.1-70B it attains 3.85 perplexity and 72.46% average zero-shot accuracy (0.22 percentage points below full precision), and in a Llama-3.

AI-generated editorial illustration: PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers

Interpretation

The paper identifies the group-constant directions of a grouped asymmetric quantizer as a range-null subspace: a level shared within a group does not raise that group's extremes and is carried by the affine offset the format already stores, so rotation design becomes a spectral problem of steering dominant activation energy into that subspace. Prior rotation-based quantization (fixed Hadamard in QuaRot, learned rotations in SpinQuant, structured block rotations in DuQuant, and others) mainly optimizes activation geometry and treats the quantizer as given; a second line first locates outlier directions from token statistics and handles them specially. PrismQuant reverses the roles, treating the quantizer's existing free subspace as the target, with outliers entering only as energy the rotation captures rather than as directions to be found first. The paper defines the subspace and its projection decomposition (Equations 2 and 3) and notes this makes explicit a property of the existing quantizer without requiring an extra mean-subtraction operation; a four-feature numerical illustration shows the dominant component's contribution to range falling from 20 to zero after mapping.

Rotation design is cast as a Ky Fan trace maximization with a closed-form solution provably optimal for this alignment objective, realized through compact Householder reflections in compact-WY form without gradient-based training, foldable into adjacent weights or executed online as a block Hadamard plus a rank-r correction. Unlike methods that learn rotations or optimize scaling per layer, the optimal subspace here follows directly from calibration statistics, and estimation and realization are separate steps, neither requiring task-loss optimization. Proposition 1 gives the optimality proof via the Ky Fan maximum principle (Appendix B.2); Appendix B.3 derives the Householder construction and compact-WY representation and notes the online correction has rank at most r with per-token computation and storage far below a dense transform; Appendix A.1 reports that before any quantizer is attached every rotation reconstructs below a stated relative round-trip error and the folded model reproduces its bf16 reference per token.

The paper derives a range law connecting unaligned activation energy and group size to the quantization step, and uses it to discuss the metadata budget: finer groups buy both finer scale resolution and a larger alignable subspace, and at a shared activation bit budget spending on finer groups is more effective than adding affine directions within larger groups. This extends group size from a question of metadata cost and local resolution into a design variable that also sets the dimension of the alignable subspace, with metadata-matched ablations demonstrating the trade-off directly. Appendix C derives a deterministic bound of residual energy on aggregate squared range and a two-factor range law whose only approximation is a single measured crest-factor ratio (near one at the q/k/v input and within a stated range at the down-projection input); Table 6a compares group sizes and represented directions at a shared bit budget and reports the default configuration as better on both Llama models; Appendix C.3 reports that the predicted step ordering agrees with measured perplexity ordering (a Spearman correlation is given).

In W4A4KV4 post-training quantization experiments, PrismQuant achieves the lowest perplexity among the compared methods across the Llama, Qwen, and Mistral families and extends to sparsely routed mixture-of-experts models, while the deployment study shows small online transform overhead. Against the metadata-matched Hadamard baseline, the paper reports closing stated fractions of the bf16 perplexity gap on 3B, 8B, and 70B, improving all eight tasks over Hadamard on 70B, and on Qwen3-30B-A3B, with one rotation per expert and the router untouched in bf16, recovering most of the accuracy Hadamard loses. Table 1 reports perplexity and eight-task zero-shot accuracy for Llama-3.2-3B, Llama-3.1-8B, and Llama-3.1-70B, averaged over three random seeds; Table 2 reports the Qwen3-30B-A3B dense and MoE comparison; Appendices A.2 and A.3 give the Qwen3 dense family and Mistral-7B-v0.3 results; Appendix D reports prefill, decode, and memory data on a single NVIDIA A40 for Llama-3.2-3B and Llama-3.1-8B and states that the performance benchmark uses random weights, making it a kernel-swap performance benchmark rather than a model-quality evaluation.

Perspective

The result targets post-training quantization deployments that use grouped asymmetric INT4 activation quantization, GPTQ INT4 weights, and a KIVI-style KV-cache policy, and it applies to dense models from 0.6B to 70B in the Llama, Qwen, and Mistral families as well as sparsely routed models such as Qwen3-30B-A3B; rotations are applied at standard sites (residual stream, after the value projection, and online before the down projection) and can be folded into adjacent weights or executed online as a block Hadamard plus a rank-r correction. For a reader, this offers a reusable design question: first ask which directions the quantizer already represents for free, then decide where the rotation should send energy; it also offers a metadata-budget view in which group size is simultaneously a resolution lever and an alignment-capacity lever. The paper further notes that a native INT4 GEMM consuming group-wise asymmetric activations would turn the measured transform overhead into a checkpoint-level system benefit.

The range-law predictor is accurate at the q/k/v input but systematically conservative at the down-projection input, which the paper attributes to the assumption that the residual keeps its shape as aligned energy is removed, and it reports the measured crest-factor ratio as a diagnostic; this makes the predictor better suited to ordering configurations than to absolute step estimation. The optimal alignment rank varies by model and recovery is not monotone at intermediate ranks, so the paper adopts r=8 as a deployment operating point rather than a universal optimum. The deployment evidence holds within a shared packed-INT4 pipeline, and the performance backend uses per-token symmetric activation quantization and KV interfaces distinct from the paper's native group-128 asymmetric accuracy protocol, so the two do not constitute a checkpoint-level joint accuracy-speed measurement; the paper also records large-accumulator FP16-conversion failures in the shared backend as an explicit numerical limitation. In addition, the paper states that quantization may preserve or alter the underlying models' biases and unsafe behaviors, and that improved quantization accuracy should not be read as evidence of safety or fairness. For this reading, the main text and appendices provide the method, proofs, ablations, and deployment data, but some figures are presented as images, so point-by-point numerical verification would still require consulting the original figures.

Sources