DRQ minimizes worst-case reconstruction loss on a fixed quantization grid, improving perplexity and downstream accuracy for LLMs quantized by six PTQ methods without adding inference overhead
Synopsis
The authors show that lower reconstruction loss on calibration data in weight-only PTQ does not guarantee better model performance and can even degrade it on the same calibration data; they propose DRQ, which minimizes worst-case reconstruction loss over a constrained set of input activation distributions by updating only the integer codes of quantized weights while keeping bit width, grouping, scales, and zero points fixed, improving perplexity and average task accuracy across six PTQ methods and both dense and MoE models without adding inference overhead.
Figure 1: Lower reconstruction loss need not improve quantized model performance. Results are from Llama-3.2-1B-Instruct with 3-bit quantization. (a) Among 32 modifications to AWQ-quantized weights, 18 reduce reconstruction loss but increase model PPL on the same WT2 calibration data. (b) Among 560 modified weight matrices per method that reduce reconstruction loss on calibration data, 413 (73.8%) for AWQ and 264 (47.1%) for GPTQ increase model PPL on held-out WT2 data. (c) Further minimizing reconstruction loss increases held-out C4 PPL by 0.69%, whereas DRQ accepts slightly higher reconstruction loss on calibration data and lowers C4 PPL by 0.32%. Each evaluation modifies one linear layer, keeping all other weights fixed. In (a,c), reconstruction loss is divided by its initial value, and PPL changes are relative to the initial quantized model.
arXivInterpretation
The paper reports a counterintuitive phenomenon: further reducing reconstruction loss on calibration data does not necessarily improve quantized model performance and can even worsen it on the same calibration data. Prior weight-only PTQ work largely treats minimizing calibration reconstruction loss as the main criterion for preserving low-precision quality; this work isolates and characterizes the gap between calibration fit and deployment performance. On Llama-3.2-1B-Instruct with 3-bit AWQ, 32 modified weight matrices were produced across 16 layers using two WT2 calibration sets, and in 18 of 32 cases lower reconstruction loss came with higher PPL on the same calibration data; in a separate experiment, 413 of 560 replacements (73.8%) with lower reconstruction loss increased PPL.
The paper derives conditions under which changes in the input activation distribution reverse the reconstruction-loss comparison, and builds a worst-case reconstruction objective from that analysis. The analysis reduces the dependence of reconstruction loss on inputs to changes in the activation second-moment matrix and gives an exact criterion under positive semidefinite perturbations and a distribution-change budget, turning robustness into an optimizable objective rather than a heuristic regularizer. Theorem 3.1 states an if-and-only-if condition, and Theorem 3.2 reduces the inner maximization to a worst-case directional error along a single input direction obtained from an eigenvalue problem; the appendix provides proofs and a distribution change attaining this loss.
DRQ is a post-hoc refinement framework that updates only the integer codes of quantized weights while preserving the quantization format and inference operators, so it adds no inference overhead. Unlike PTQ methods that change scaling, clipping, rotations, or the reconstruction objective, DRQ improves the weight-selection criterion within the existing quantization grid and can be layered on top of multiple base PTQ methods. In vLLM benchmarks on Llama-3.1-8B W4A16 with group size 128, prefill and decoding times change within a small range after refinement and model tensors occupy the same memory; offline refinement averaged about 6.01 and 5.96 minutes with GPTQ and AWQ across 20 configurations.
DRQ improves both perplexity and average task accuracy across model sizes, bit widths, base PTQ methods, and dense and MoE architectures. The paper extends refinement benefits from a single method to six representative PTQ methods (including AWQ, GPTQ, and ParoQuant) and to dense and MoE models, showing gains even when the initial quantized model is already strong. For Llama-3.1-70B with GPTQ W2, C4 PPL falls from 31.77 to 27.11 and accuracy rises from 48.23% to 54.56%; for Llama-3.2-3B with GPTQ W3, C4 PPL falls from 30.34 to 22.11 and accuracy rises from 55.83% to 60.24%; for Qwen3-30B-A3B at W3, WT2 and C4 PPL decrease and accuracy improves on five of seven tasks.
Perspective
The work targets post-hoc refinement for weight-only post-training quantization, suited to deployment pipelines that already have quantized weights from a base PTQ method and can obtain calibration activations; DRQ updates integer codes within the existing quantization grid while keeping bit width, grouping, scales, zero points, and inference operators fixed, making it relevant to researchers and engineering teams who want to improve low-bit model quality without changing the inference stack or adding inference overhead. The paper validates on Llama-3.2-Instruct (1B, 3B), Llama-3.1-Instruct (8B, 70B), and Qwen3-30B-A3B under W4A16, W3A16, and W2A16 with group size 128, covering GPTQ, AWQ, RTN, OmniQuant, GPTAQ, ParoQuant, and QuaRot combined with GPTQ.
Several open questions remain for a careful reader: the worst-case direction depends on the weight difference and the activation second-moment matrix, and changing integer codes changes that direction, so DRQ re-estimates it each round and how estimation accuracy interacts with the acceptance rule deserves further study; the paper reports that severe AWQ W2 degradation persists, indicating a scope to how much quality refinement can recover; the sensitivity sweep over the parameters shows that a stronger worst-case penalty does not consistently improve performance and that favorable settings depend on the base PTQ method; in MoE models some experts receive very few or zero calibration tokens, and those layers use RTN initialization or fall back to a weight-distance term, so their refinement behavior differs from dense layers.
