OnlineQAT uses on-policy distillation on student-generated prefixes to lift Qwen3-1.7B to a 57.28 six-task average at W3A16, 2.90 points above ReasoningQAT
Synopsis
OnlineQAT is a two-stage ultra-low-bit recovery framework that first obtains a usable W2/W3 initialization through block-wise quantization-aware training, then performs on-policy distillation on responses sampled from the quantized student, with a frozen full-precision teacher supplying a sampled reverse-KL signal at each visited prefix; on Qwen3-1.7B it reaches six-task averages of 57.28 at W3A16 and 32.52 at W2A16, improving over ReasoningQAT by 2.90 and 0.44 points, suggesting that student-visited states offer a recovery signal beyond fixed-completion training, particularly at three bits.
Interpretation
The paper identifies a prefix-distribution mismatch between offline recovery and deployment: existing recovery stages optimize on fixed completions or teacher-generated answers, whereas the deployed quantized model conditions on tokens it generated itself, so quantization errors can move it into states absent from offline recovery data. Relative to BitDistiller, EfficientQAT, and ReasoningQAT, which evaluate recovery on fixed calibration sequences or reference completions, OnlineQAT moves end-to-end recovery onto the prefixes the student actually visits while retaining block-wise QAT initialization. The mismatch is argued through the schematic contrast in Figure 1 and the reasoning in the introduction, making this a methodological motivation rather than an independently measured effect.
OnlineQAT obtains the best six-task average among the compared quantized methods on Qwen3-1.7B: 57.28 at W3A16 and 32.52 at W2A16. Against the strongest offline baseline, ReasoningQAT, it improves by 2.90 points at W3A16 and 0.44 points at W2A16; against the fixed-prefix distillation baseline KD it improves by 3.62 points at W3A16. Based on the unweighted six-task mean in Table 1, with all recovery methods starting from the same Stage 1 checkpoint within each bit width and Stage 2 runs using AdamW on eight GPUs with effective batch 64.
The W3A16 gain is broad, while the W2A16 gain is smaller and task-dependent. Relative to ReasoningQAT, W3A16 improves GSM8K, MMLU-R, MATH-500, LiveCodeBench, and GPQA-Diamond, with the largest changes on LiveCodeBench and GPQA-Diamond, while IFEval decreases by 0.8 points; W2A16 gains on GSM8K, MMLU-R, and GPQA-Diamond but trails on MATH-500, IFEval, and LiveCodeBench. Per-task scores come from Table 1; the paper itself describes the W2A16 improvement as modest and task-dependent.
Methodologically, Stage 2 uses a one-sample policy-gradient estimator of sampled reverse KL, applies loss only to response tokens, includes no cross-entropy term, and freezes the teacher and zero-points during Stage 2. Unlike ReasoningQAT, which uses forward KL on gold completions and retains the teacher's top-20 logits, OnlineQAT samples from 32,757 OpenThoughts prompts at temperature 0.6 and trains W3A16 for 40 steps and W2A16 for 30 warm-up steps followed by 120 steps. The objective and estimator are given in Equations 2 through 4, and hyperparameters and step counts are listed in Section 4.1; the paper notes that the constant term in the reverse-KL policy gradient is omitted because its expectation is zero.
Perspective
The result targets engineering and research settings that need to deploy reasoning models such as Qwen3-1.7B at W3A16 or W2A16, especially teams that want to replace fixed-completion recovery with on-policy distillation without changing model architecture. The method depends on a usable Stage 1 checkpoint, so block-wise QAT initialization remains a prerequisite; Stage 2 samples responses from the student and therefore incurs rollout generation cost. The paper frames its findings as initial results and explicitly motivates broader evaluation of on-policy recovery for ultra-low-bit language models.
Relative to ReasoningQAT, OnlineQAT changes both the prefix distribution and the KL direction, so the comparison does not isolate either factor; the paper also notes the absence of an offline reverse-KL control. The study covers one model without multiple training seeds, and the W2A16 gain is task-dependent, so the stability of the average difference still needs broader evaluation. In addition, OPD incurs rollout generation, and the paper states that Table 1 supports an empirical comparison rather than a causal or compute-efficiency claim. Readers interested in the temperature-scaling details in the equations or per-task variance should consult the original formulas and tables.
