Skip to main content
Back to timeline
arXivSource publication:

OSFP4 jointly optimizes diagonal smoothing and block scales, reaching 77.03 average W4A4 accuracy on Llama-3.1-8B-Instruct while retaining about 94-97% of vendor NVFP4 prefill throughput

Synopsis

The work introduces OSFP4: by analyzing a multiplicatively dithered FP4 quantizer it obtains a differentiable approximation of the matrix-product quantization error, then jointly optimizes a diagonal smoothing matrix and weight block scales per linear projection, with separate losses for RTN and GPTQ-style SIC rounding; on Llama-3.1-8B-Instruct the W4A4 configuration attains the highest average accuracy among the evaluated methods while retaining roughly 94-97% of vendor NVFP4 prefill throughput.

Source-provided article image: OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization
Figure 1 ·

Figure 1: Normalized mean-squared error of randomized FP4 quantization, ϕ ⁡ ( x ) = 𝔼 ⁡ [ ( x − Q ~ FP4 ​ ( x ) ) 2 ] x 2 \phi(x)=\frac{\mathbb{E}[(x-\tilde{Q}_{\mathrm{FP4}}(x))^{2}]}{x^{2}} , and its deterministic counterpart, ϕ exact ​ ( x ) = ( x − Q FP4 ​ ( x ) ) 2 x 2 \phi_{\mathrm{exact}}(x)=\frac{(x-Q_{\mathrm{FP4}}(x))^{2}}{x^{2}} .

arXiv

Interpretation

The paper derives an approximation of the expected squared Frobenius norm of the matrix-product error under diagonally smoothed NVFP4 quantization, as a function of smoothing parameters and group scales, distinguishing RTN from SIC rounding. SmoothQuant-style diagonal scaling was aimed at balancing infinity norms for INT constellations, whereas the small E2M1 dynamic range of NVFP4 calls for a different criterion; this expression folds the rounding procedure actually used into the smoothing optimization. The derivation builds on a multiplicatively dithered FP4 quantizer and a high-resolution quantization assumption, with full derivations in Appendix I; Lemma 1 gives the minimum of the dithered quantizer's relative MSE and the interval on which it is attained.

Based on that approximation, OSFP4 runs a two-stage pipeline: continuous joint optimization of diagonal smoothing and block scales, followed by discrete E4M3 weight-scale selection, with absmax-to-6 as the default activation block scaling. Compared with scale-only methods such as H-Scale and ScaleSweep or rotation-based MR-GPTQ, this approach optimizes smoothing and scaling under one objective, and the diagonal transform can be folded into LayerNorm/RMSNorm. Appendix B details the Adam iterations, log-parameterization, lookup-table interpolation gradients, and scale normalization; Appendix C gives distinct E4M3 scale-selection algorithms for RTN and SIC.

On Llama-3.1-8B-Instruct, OSFP4 W-SIC/X-RTN reaches 77.03 average W4A4 accuracy, above FP-Quant GPTQ at 76.34, MR-GPTQ at 76.18, and H-Scale at 76.67; W4A16 OSFP4 W-SIC reaches 78.42, above H-Scale at 78.00 and below the BF16 baseline of 79.22. Under a common evaluation protocol, a diagonal transform alone yields higher average W4A4 accuracy than rotation-based MR-GPTQ; ablations show the lowest perplexity when joint smoothing, weighted-MSE scale selection, and SIC are combined. Evaluation covers MMLU-CoT, GSM8K, HellaSwag, and WinoGrande; locally prepared models share one BF16 checkpoint and a calibration set of 1024 sequences of 2048 tokens from FineWeb-Edu, served through vLLM on an RTX 5090; Appendix H reports relative MatMul MSE across seven projections and WikiText-2 perplexity ablations.

On an NVIDIA B200, OSFP4 retains about 94-97% of vendor NVFP4 prefill throughput, and a control that sets the smoothing matrix to identity achieves essentially the same throughput as OSFP4. This indicates the gap to the vendor NVFP4 path mainly reflects the runtime implementation of explicit diagonal smoothing before the attention output and MLP down projections, where it cannot be absorbed into a preceding LayerNorm, rather than the values of the smoothing coefficients. Appendix G reports the full prefill and decode workload sweep, a timing protocol of three independent repetitions of 30 inference runs each, and comparisons against BF16, vendor NVFP4, and unit-smoothing NVFP4.

Perspective

The results target post-training quantization deployment of LLMs using NVFP4 (E2M1 values, one E4M3 block scale per 16 input channels, FP32 tensor scale), with primary evidence from Llama-3.1-8B-Instruct and additional results in the appendix for Llama-3-8B, Qwen3-8B, Qwen3-30B-A3B-Instruct, and Gemma-4-31B-IT. Diagonal smoothing can be folded into LayerNorm/RMSNorm where present, but must be applied explicitly at runtime before the attention output and MLP down projections, so throughput gains are workload dependent. The default activation block scaling is absmax-to-6, and the method is complementary to online activation-scale search (ScaleSearch/ScaleSweep); for MoE models the authors use expert-specific smoothing matrices, but because a route-aware serving path was not implemented, only W4A16 results are reported for Qwen3-30B-A3B-Instruct.

The zero-mean, unit-variance, and cross-entry/cross-block independence assumptions in the dithered-quantizer error model are modeling assumptions rather than exact properties of quantization error, and the SIC analysis additionally relies on a high-resolution quantization assumption. The projection-level measurements in Appendix E use the same empirical activation population for optimization and measurement, which the authors describe as a calibration diagnostic rather than a held-out generalization test. Several competitors (ScaleSweepMSE, 4over6, SOAR, H-Scale and its W4A4 extension) are author reimplementations, and the released NVIDIA checkpoint's preparation is outside the controlled procedure, so cross-method comparisons should be read with those provenance differences in mind. In addition, some numeric values are not fully rendered in the loaded text, so exact figures should be taken from the original tables.

Sources