Skip to main content
Back to timeline
arXivSource publication:

An equivalent reparameterization of the output head cuts Phi-4-mini W4 AW-MSE KL from 0.936 to 0.256 while keeping a 10.8% batch-one latency gain

Synopsis

The work proposes softmax reparameterization: before quantization, subtract a scalar multiple of the vocabulary-row mean from every output-head row and pick the coefficient by validation KL separately for RTN, AW-MSE and full-Hessian GPTQ, preserving the full-precision softmax distribution and the trained decoder while improving low-bit output-head fidelity, for example lowering Phi-4-mini W4 AW-MSE KL from 0.936 to 0.256 and reducing batch-one generation latency by 10.8% relative to a BF16-head baseline in packed W4 deployment.

AI-generated editorial illustration: Softmax Reparameterization for Output-Head Quantization

Interpretation

The authors introduce a post-training one-dimensional search inside the additive softmax equivalence class of the output head, subtracting a scalar multiple of the vocabulary-row mean and selecting the coefficient by the quantized model's validation KL to the source, with the search including both the original head and fixed mean-centering. Prior work used the softmax symmetry and mean-centering for attention keys or pretraining stability; here the freedom is used to choose a pre-quantization output-head representative by post-quantization predictive fidelity rather than by weight norm. The paper derives the equivalence (softmax depends only on relative logits) and evaluates seven heads with a split of 128 WikiText fitting articles, 16 validation and 16 evaluation articles; Phi's 14-point GPTQ sweep takes about twelve minutes.

At W4 the gains concentrate on heads whose baseline quantization substantially distorts predictions: Phi-4-mini AW-MSE KL falls from 0.936 to 0.256, XGLM RTN KL from 2.129 to 0.143, and BLOOM/BLOOMZ also improve, while Gemma 3/4 and Qwen3.5 already have low baseline KL and change little. Relative to fixed mean-centering, the validation-selected coefficient is better on Phi, BLOOM and BLOOMZ; on XGLM all three quantizers select mean-centering itself. Table 2 reports raw, centered and reparameterized KL for four heads under RTN/AW-MSE/GPTQ; on Phi the validation-selected coefficient lowers RTN test KL from about 1.23 to about 0.35, roughly half of fixed centering.

The gains survive stronger GPTQ calibration, exact per-channel scaling and affine quantization, and transfer to C4 and OpenWebMath: the frozen WikiText-selected coefficient beats mean-centering in all 18 comparisons where it differs from 1 and ties in the remaining six XGLM cases. This indicates the effect is not an artifact of weak calibration or missing scaling, but a step complementary to existing quantization improvements. On Phi, raising GPTQ calibration from 1,024 to 65,536 states lowers raw KL from 0.160 to 0.089 and reparameterization further to 0.035; scaled AW-MSE falls from 0.240 to 0.172 and affine RTN from 0.750 to 0.259, with paired bootstrap intervals excluding zero.

Mechanistically, the selected representative can increase ordinary logit reconstruction error while moving that error into directions softmax barely responds to: on Phi, going from coefficient 0 to 4 raises raw logit-error energy by factors of 3.0 under RTN and 2.7 under AW-MSE while Fisher-weighted error falls by 72-73%. This explains why selecting a representative by weight or logit reconstruction error fails: in Appendix E.2 both weight MSE and projected MSE select coefficient 1 in all nine comparisons, while the diagnostic-KL optimum differs every time. Table 4 reports actual KL, the Fisher quadratic and probability-bin diagnostics on 8,176 held-out states; under AW-MSE the absolute error energy on high-probability entries falls by more than half while the error share on low-probability entries rises.

Perspective

The method targets post-training quantization where only the output head is quantized and the decoder is frozen, and applies to linear-softmax output heads; nonlinear logit paths such as tanh soft-capping need a rank-one correction costing one hidden-dimension dot product and one vocabulary broadcast. It suits deployers already using packed W4 kernels such as Marlin/vLLM, adding no inference operation for shift-compatible heads; for tied-weight models the BF16 input embedding is retained and a separate packed output head is added, raising Phi's resident weights from 7.17 to 7.47 GiB. Coefficients are selected per quantizer, and frozen WikiText coefficients still beat mean-centering on C4 and OpenWebMath.

Coefficient choice depends on the calibration distribution: Appendix C.3 shows larger C4 and OpenWebMath validation sets can select coefficients different from WikiText, and in the five-language mC4 study French RTN still disagrees across ten seeds at 1,024 documents. In the FLORES multilingual diagnostic BLOOM KL rises for Arabic and Hindi, so frozen coefficients are not universally transferable. W2 is framed as a compression stress test, and absolute quality remains poor in several configurations. The mechanism analysis centers on Phi at W4, and the authors note these measurements explain Phi's KL reduction without establishing which other heads will benefit. The loaded text is the full paper with appendices, but some table cells appear as placeholders, so specific cell values should be read from the numbers stated explicitly in the prose.

Sources