Skip to main content
Back to timeline
arXivSource publication:

ISTA-DASLab proposes disaggregated quantization: separate formats for prefill and decode, more than doubling 1-bit decode accuracy

Synopsis

The work proposes disaggregated quantization (DQ), which specializes computation formats, weights and storage placement to the prefill and decode phases of LLM inference, together with the QADD training method; on Qwen 3 and Gemma 3, removing activation quantization only on decode improves accuracy on decode-heavy tasks without increasing inference cost, training separate NVFP4 prefill weights accelerates prefill while raising accuracy at 2–3-bit decode, and on Qwen3.8-27B 1-bit GGUF decoders it more than doubles MMLU-Pro and MMMU-Pro accuracy, with offloaded disaggregated prefill streaming prefill weights from SSD delivering a time-to-first-token speedup over the weight-only baseline at 8K context.

AI-generated editorial illustration: Disaggregated Quantization: Specializing LLM Prefill and Decode

Interpretation

The paper introduces quantization-aware distillation with disaggregation (QADD), which uses the SFT label mask to propagate the prefill/decode split into quantized layers: prompt (user turn) uses the prefill pathway while response (assistant turn) uses decode, and both pathways are trained toward a common response objective in one forward-backward pass, supporting shared master weights or prefill-only adaptation to a frozen decoder. Prior quantization-aware distillation did not separate the computational pathways of the two inference phases; QADD moves phase specialization from deployment down into training, letting one model hold two phase-specific representations. Training minimizes against a frozen BF16 teacher on 100M tokens of the Tülu 3 SFT corpus with sequence length 2048, global batch size 64 and 2450 steps, FP32 master weights with a straight-through estimator, across Qwen 3 at 0.6B/1.7B/4B/8B and Gemma 3 at 1B/4B/12B.

The paper presents three complementary schemes: format disaggregation (hardware-native quantized prefill with low-bitwidth weight-only decode), full disaggregation (additionally training native NVFP4 prefill weights), and offloaded disaggregated prefill (ODP, which streams prefill weights from SSD block by block and reuses device buffers). Existing work mostly varies phase-specific weight precision or accelerates prefill alone; DQ unifies hardware-native prefill with compressed weight-only decode and further brings the residency of prefill weights into the design. Table 1 reports prefill and decode speedups, device memory and both accuracy families for BF16, NVFP4A16, NVFP4, 3-bit and 2-bit weight-only and their disaggregated variants on Qwen 3 and Gemma 3; ODP overlaps loading with compute using two device block buffers and carves buffer space from decode weights on DGX Spark.

On decode-heavy tasks, disabling activation quantization only on decode improves accuracy without increasing weight storage or inference cost, while on prefill-heavy tasks this change has little effect. This shows quantization error is asymmetric across the two phases and that the gain from phase separation lands mainly on the decode side, supporting format disaggregation as a plug-in replacement. The paper first quantizes each phase to NVFP4 in isolation: decode-only quantization incurs a fraction of the accuracy loss of prefill-only quantization on most models, with the largest gap on Gemma3-1B, while on prefill-heavy tasks prefill-only quantization incurs a fraction of the loss of decode-only quantization on seven of eight models. In Table 1, format disaggregation is 2–3% faster than the non-disaggregated scheme on decode, with prefill speedups up to about 1.49x on Qwen3-8B and 1.67x on Gemma3-12B.

Training an NVFP4 prefiller augments arbitrary frozen weight-only checkpoints: on Qwen3.8-27B 1-bit GGUF decoders it more than doubles MMLU-Pro and MMMU-Pro accuracy over weight-only inference while making prefill faster and keeping device memory unchanged via ODP. This lets disaggregated quantization reuse existing quantized decoders, including ones produced by closed-source or complex algorithms, by training only a prefill model, so the decode quantization pipeline can remain a black box. Eight publicly released Unsloth GGUF checkpoints, including vector-quantized formats such as IQ2_XXS, are used with one complete evaluation per format; gains are largest at 1-bit decode, smaller at 2-bit, and at 3-bit the NVFP4 prefiller leads to slight degradation; at 8K context time to first token falls from several seconds to several seconds, with a speedup range across 4K–32K contexts and ODP slower at shorter prompts due to SSD loading.

Perspective

The work targets post-trained instruction-following models, since disaggregated quantization needs a logical input/output separation; evaluation covers decode-heavy reasoning tasks and single-turn prefill-heavy tasks, with batch-one decode, prefill-stack timings on DGX Spark, and ODP time-to-first-token measurements in llama.cpp. Format disaggregation suits single-user serving where device memory is scarce and the original model's interactivity must be preserved; full disaggregation suits low-concurrency disaggregated serving, where two checkpoints are resident on accelerators anyway, and local single-user serving via ODP; the prefiller scheme suits settings where an existing weight-only quantized checkpoint should be reused without retraining the decoder. The authors release a reproduction codebase, a llama.cpp fork with ODP support, and the trained Qwen3.8-27B prefillers.

The paper does not evaluate highly batched performance or multi-turn and agentic behavior; the authors note that in multi-turn use cached assistant tokens retain decode-produced KV entries, while rebuilding their cache through prefill can produce different representations for the same token history, and robustness to this cache-policy dependence remains untested. ODP is slower at shorter prompts because of SSD loading, and the authors consider the idea not to transfer seamlessly to mixture-of-experts models. Prefillers change the representations that condition generation, and the appendix shows that on MMLU-Pro median generation length falls for most formats while mean length rises, so unchanged decode weights do not by themselves guarantee unchanged generation cost. Interoperability between prefillers and decoders holds only partially at the lowest bitwidths, and which decoder a prefiller was trained with affects the gain. This summary is based on the full text, but some numeric values appear as placeholders in the text, so exact point values should be checked against the original tables.

Sources