ResidualQuant anchors on the final-loop KV and stores 2-bit residuals, cutting looped-Transformer KV storage by 80.7% at near-BF16 accuracy under INT2/4
Synopsis
To address the KV cache memory bottleneck that grows with recurrent depth in looped Transformers, the authors propose ResidualQuant, which uses the final-loop KV states as an anchor and represents the remaining loops as low-precision residuals, combined with least-squares scaling, rotations applied to the residuals, and loop-wise INT4/INT2 mixed precision; across Ouro-1.4B and Huginn-3.5B on mathematical reasoning and code generation benchmarks it improves the accuracy-memory tradeoff over rotation-based KV quantization, reducing theoretical KV storage by 80.7% relative to BF16 while retaining accuracy close to BF16 under mixed precision.
Figure 1 : ResidualQuant progressively recovers BF16-level accuracy while reducing KV storage and improving decode throughput. Results on Ouro-1.4B with group size g = 32 g=32 in every quantized loop. Left: Starting from ① direct INT2 quantization, we successively introduce ② Last-loop residual quantization, ③ least-square scaling, ④ rotation, and ⑤ mixed precision to achieve ResidualQuant . Accuracy increases from 27.8% to 76.0%, matching the BF16 baseline of 75.0%, while KV cache storage is reduced by 80.7%. Right: At 8k context on an RTX 5090, ResidualQuant achieves 3.27 × 3.27\times the peak decode throughput and 4 × 4\times the largest feasible batch size compared to the BF16 counterpart.
arXivInterpretation
Introduces a loop-aware KV quantization framework that uses the final-loop KV as a shared anchor and stores the remaining loops as low-precision residuals, augmented with least-squares scaling, orthogonal rotations applied to the residuals rather than to the KV states, and loop-wise mixed precision with INT4 for the anchor and INT2 for residuals. Prior work reduced looped KV cost by sharing or reusing KV across loops (MoR, PLT, MELT), which can discard loop-specific information, while generic KV quantization degrades substantially below 4 bits. This work instead preserves each loop's distinct KV and compresses only the inter-loop differences. Evaluated on Ouro-1.4B (4 loops) and Huginn-3.5B (32 loops partitioned into eight groups of four) across GSM8K, MATH500, HumanEval, and MBPP with shared prompts and greedy decoding, compared against direct quantization and the OptR-H rotation baseline at matched KV budgets.
Under mixed INT2/4, ResidualQuant nearly matches the BF16 average score (52.37 vs. 52.94) while reducing theoretical KV storage by 80.7%; at the same configuration direct quantization and OptR-H fall 13.31 and 5.77 points below BF16. At comparable KV budgets, residual quantization without rotation achieves a higher average score than the rotation-based OptR-H baseline across all four precision configurations, reaching a 4.95-point gain under mixed INT2/4 with group size 16 (52.12 vs. 47.17). The main results table covers all model-benchmark combinations, with effective bitwidth accounting for quantized codes, FP8 scales and offsets, and the extra BF16 least-squares coefficients stored for residuals.
Ablations show the components are complementary: the Last-loop anchor raises MATH500 from 72.2% to 75.2% under INT2 and from 72.8% to 76.0% under INT2/4 relative to the Previous-loop anchor; least-squares scaling raises 70.2% to 73.8% in the INT2/4 Last-loop setting; adding rotation to residual quantization reaches 75.2% at INT2. Reference choice, scaling method, rotation method, and loop-wise precision allocation are compared within one evaluation, and shared-anchor reconstruction is shown to need only two KV cache reads while chained reconstruction grows with distance from the anchor. Ablations use Ouro-1.4B on MATH500, with a full appendix table covering Norm-ratio scaling and both OSCAR and OptR-H rotations.
KV compression translates into measured throughput gains: at 8k context and batch size 4 decode throughput rises from 111.4 to 192.5 tokens/s, and at 16k context and batch size 2 from 34.0 to 93.1 tokens/s; the smaller footprint also raises the largest feasible tested batch from 4 to 16 at 8k (364.9 tokens/s) and from 2 to 4 at 16k (141.3 tokens/s). Beyond fixed-batch speedups, the work quantifies the batch-capacity gain from compression and provides a vLLM integration with custom CUDA/Triton kernels so reconstruction cost does not offset the saved memory traffic. Measured on an RTX 5090 with Ouro-1.4B against BF16 vLLM FlashAttention, doubling batch sizes from 1 to 128, generating 128 tokens per run, with throughput excluding prefill and averaged over three runs.
Perspective
The results target inference for looped or recursive Transformers that repeatedly apply shared blocks, especially server-side deployments where KV cache grows with recurrent depth and larger batches are needed. Evaluation covers Ouro-1.4B (4 loops) and Huginn-3.5B (32 loops partitioned into eight groups of four) on GSM8K, MATH500, HumanEval, and MBPP, with weights kept in BF16 and KV evaluated in BF16, INT4, INT2, and INT2/4. The method needs no retraining or pretrained-weight updates, and rotation matrices are calibrated offline and then fixed, so it can plug into existing inference stacks; the authors also demonstrate combination with FlashLoop's sparse update and attention mechanisms and apply the same residual reconstruction principle to W4A4 activation quantization.
Several throughput improvement percentages are missing from the loaded text, so those gains can only be judged from the concrete tokens/s and batch-size changes reported in Section 5.3; the Huginn-3.5B activation-quantization protocol differs from its KV evaluation protocol, so cross-table comparisons need care; LoopQ was reproduced without the authors' code, so its absolute numbers should be read with that context; larger loop groups (8 or 16) show similar GSM8K accuracy but appear only in the appendix, leaving behavior on other benchmarks an open question; and the gain from combining residual quantization with different rotation methods varies with precision and rotation method, so the best pairing still needs per-setting calibration.
