Skip to main content
Back to timeline
arXivSource publication:

SoloQ quantizes diffusion language models to 4 bits without calibration, holding accuracy on LLaDA, Dream, Fast-dLLM v2 and Nemotron-Labs-Diffusion while beating calibration-based baselines

Synopsis

The authors present SoloQ, which uses a K-RPBH rotation to map weights and activations into a predictable near-Gaussian basis plus a scale-only bias correction for calibration-free 4-bit quantization, and quantizes the KV cache at block commit for block-diffusion models; across LLaDA-Base, LLaDA-1.5, Dream-7B, Fast-dLLM v2 and Nemotron-Labs-Diffusion, SoloQ-C and SoloQ-N retain accuracy under W4A4 and W4A4KV4 and exceed calibration-based baselines such as DLLMQuant++, STaR-Quant and FAIR-Calib on average, with the NVFP4 path also cutting peak memory and speeding up end-to-end inference.

Source-provided article image: SoloQ: Calibration-Free Quantization for Diffusion Language Models
Figure 1 ·

Figure 1: Profiling and motivation for SoloQ . (a) Memory breakdown at batch size 128 and aggregate linear/attention latency at batch size 1 for Fast-dLLM v2 ( Wu et al., 2026a ) across context lengths on an H200 GPU. (b) Layer-wise variation in down_proj input activation scale across denoising steps, normalized to the first step. (c) Accuracy of PTQ ( Ashkboos et al., 2024 ; Frantar et al., 2022 ) with WinoGrande/MMLU calibration versus calibration-free SoloQ .

arXiv

Interpretation

SoloQ rotates weights and activations into a K-RPBH basis whose coordinate marginal is predictable, enabling quantization without calibration data. Existing dLLM post-training quantization methods rely on calibration data to track activation distributions that shift across masking states and denoising steps; SoloQ instead transforms the distribution into a predictable form and restores the component along the original vector with a scale-only bias correction. The paper proves cross-block variance averaging in Proposition 1 and validates it on FFN down-projection activations from Dream-7B and Nemotron-Labs-Diffusion; the ablation shows that combining the randomized permutation and cross-block mixer lowers KL from 0.11396 to 0.00413 and block-energy imbalance from 67.89% to 22.12%.

The same predictable rotated representation supports both distribution-matched codebooks (SoloQ-C) and hardware-native NVFP4 (SoloQ-N). Prior calibration-free quantization largely targeted image or video diffusion transformers, or covered weights only; SoloQ extends the idea to diffusion language models and lets both quantization paths share one rotation basis. On LLaDA-Base, SoloQ-C averages 57.74 and SoloQ-N 57.60, both above DLLMQuant++ at 54.29; on LLaDA-1.5, SoloQ-N reaches 68.48, a 4.17-point gain over DLLMQuant++; Appendix D.2 shows that the uniform INT4 variant SoloQ-I still averages 56.98, 67.44 and 55.32.

For block-diffusion models, SoloQ quantizes the KV cache at block commit, compressing only persistent states without perturbing the actively denoised block. Existing dLLM quantization work mainly targets full-sequence diffusion models; SoloQ extends the calibration-free design to persistent KV states in block-diffusion models, using KIVI-style per-channel asymmetric quantization for keys and the SoloQ activation path for values. On Fast-dLLM v2 and Nemotron-Labs-Diffusion, RTN and AWQ collapse under W4A4 (averages 6.67, 6.61 and 1.30, 1.28), and QuaRot is not applicable to Fast-dLLM v2 because its rotation cannot support the model dimensions, whereas SoloQ maintains strong accuracy under W4A4 and W4A4KV4; in the roughly 3 BPE KV comparison, SoloQ-C reaches 53.31 and SoloQ-N 53.53, above KIVI 50.26, OScar 46.62 and TurboQuant 4.03.

K-RPBH is far faster than dense Haar rotation while preserving accuracy, and improves on RPBH. RPBH's blockwise structure leaves residual inter-block energy imbalance and lacks cross-block correction; K-RPBH adds a Haar-random orthogonal cross-block mixer that redistributes the residual block statistics. Rotation latency drops from 98.135 ms for Haar to 5.779 ms on Fast-dLLM v2, and from 0.559 ms to 0.048 ms on LLaDA-1.5; relative to RPBH, K-RPBH gives 73.86/64.70 HellaSwag/MMLU on LLaDA-1.5 versus 73.08/63.48.

Perspective

The results target researchers and engineers who need to deploy diffusion language models under limited memory or on edge devices, and apply to W4A4 and W4A4KV4 settings for full-sequence dLLMs (LLaDA, Dream) and block-diffusion dLLMs (Fast-dLLM v2, Nemotron-Labs-Diffusion). SoloQ-C uses 8-bit Tensor Core kernels while SoloQ-N uses native NVFP4 kernels; the paper also implements an NVFP4 accelerator on a Zynq UltraScale+ ZCU104 FPGA, reporting end-to-end speedup and higher energy efficiency than BF16. Two natural next steps are extending the predictable rotated representation to more diffusion language models and lower bit-widths, and folding the rotation into adjacent weights to cut online overhead.

Online activation rotation adds inference overhead; folding it into adjacent weights reduces cost, but Appendix D.3 shows the folded topology can bring model-dependent accuracy degradation under low-bit quantization, especially for block-diffusion models. SoloQ-C's non-uniform codebooks represent centroids losslessly in INT8, so they cannot directly exploit native 4-bit Tensor Core execution and yield lower speedups than SoloQ-N. The full latency benefit of SoloQ-N depends on hardware support for native NVFP4 computation. In addition, preserving the attention-sink position in full precision has little effect under W4A4KV4, but becomes beneficial under the more aggressive W4A4KV2, so the boundary of ultra-low-bit KV quantization remains worth watching.

Sources