G²PTQ refreshes gradients and Hessians before quantizing each Transformer block, halving KL divergence versus GPTAQ and adding 6.45% QA accuracy at 2-bit
Synopsis
The work introduces G²PTQ, a post-training quantization framework that combines first- and second-order information under a globally supervised, block-wise objective: it recomputes gradient and Hessian estimates before quantizing each Transformer block to avoid staleness and uses trust-region scaling to bound the exact first-order compensation step, achieving lower KL divergence and higher downstream QA accuracy than baselines such as GPTQ, GuidedQuant, and GPTAQ across 13 dense and 2 MoE models from 0.6B to 125B parameters at 2/3/4-bit weight and 4-bit weight-activation settings.
Interpretation
G²PTQ unifies GPTQ-style second-order compensation with first-order gradient compensation under a block-wise supervised objective: intermediate blocks minimize block-wise MSE between quantized and full-precision block outputs, while the final block uses KL divergence, retaining global supervision while adding first-order information. Prior global-objective methods such as GuidedQuant use only second-order information, while methods that reintroduce the first-order term such as FOEM remain confined to a layer-wise MSE objective; this work merges both into one block-wise objective. The paper provides Hessian approximation theorems for both block-wise MSE and KL divergence (Theorems 1 and 2) and argues that block-wise MSE captures nonlinear interactions within a block while final-block KL directly targets the predictive distribution.
Refreshing gradient and Hessian estimates before quantizing each Transformer block avoids the staleness that affects global-objective methods with fixed Hessians as quantization proceeds. Existing global-objective methods (OAC, GuidedQuant, YAQA) compute the Hessian once via a full-model backward pass and keep it fixed; G²PTQ instead performs one backward pass through each block alone to refresh the estimates. Ablation on Qwen3-0.6B at 3-bit shows block-wise alignment raises average QA accuracy from GPTQ's 33.09% to 37.40%, and refreshing the information lifts it further to 38.61%; the paper also notes refreshing is a prerequisite for gradient compensation, since gradients are zero at the initial full-precision weights.
To address exploding weight updates caused by the numerical scales of gradients and Hessians under exact first-order compensation, a trust-region scaling mechanism dynamically bounds the gradient step. FOEM approximates gradients with a first-order Taylor expansion for numerical stability, but that approximation is exact only for layer-wise MSE; G²PTQ instead computes exact gradients via backpropagation and uses trust-region scaling to keep updates within the neighborhood where the Taylor approximation is valid. In the ablation table, removing trust-region scaling raises KL to 2.53e+01, PPL to 1.29e+12, and drops QA accuracy to 28.97%, which the paper uses to argue the mechanism is necessary for stability.
The paper derives efficient implementations of block-wise Hessian approximation and exact gradient compensation, keeping complexity comparable to vanilla GPTQ, and reports practical calibration costs. Recursive updates and a lazy-batch strategy reduce the naive implementation's high complexity to the same level as GPTQ, complemented by Triton kernels and distributed quantization optimizations for large MoE models. The paper reports quantizing an 8B model in 2.24 hours with 12.54 GB of memory on a single accelerator; the 125B Qwen3.8-Flash-Next takes 86.72 GB per device and 20.30 hours with G²PTQ* across 8 accelerators, versus 19.48 GB and 7.98 hours for GPTQ.
Perspective
The result targets deployment settings that need to compress LLM weights and activations without retraining, and applies to the model families evaluated in the paper (LLaMA3, Qwen3, Qwen3.5, Qwen3-MoE, and Qwen3.8-Flash-Next) at 2/3/4-bit weight and 4-bit weight-activation settings. It enables follow-up work to use refreshed first- and second-order information under a block-wise supervised objective, and to treat trust-region scaling as a general means of stabilizing exact gradient compensation; the paper also shows a trust-region threshold tuned for 3-bit transfers directly to 2-bit, 4-bit, and Qwen3-0.6B-Base without re-tuning. For practitioners, this means quantizing an 8B model on a single accelerator (2.24 hours, 12.54 GB) or a 125B MoE model on an 8-accelerator node is feasible, at the cost of higher calibration memory and time than GPTQ.
The reported results are for specific model families, bit-widths, and calibration data, so behavior under other architectures or data distributions still needs verification. Ablations show G²PTQ generally needs more calibration samples than GPTQ to reach its best accuracy, leaving the trade-off between calibration set size and the number of output channel groups open. Hessian-guided weight clipping helps on Qwen3 but yields no benefit or slight degradation on Qwen3.5, and the paper leaves selecting the optimal clipping method to future work. In addition, this evidence bundle is a full-text parse in which some appendix tables and figures (such as the calibration-size curve and optimal-threshold visualization) appear as references, so their specific values are not fully reproduced here and related details can only be understood from the main-text descriptions.
