Skip to main content
Back to timeline
arXivSource publication:

ByteDance Seed and UW–Madison make full-pipeline FP8 RL work: tracing entropy surges to over-clipping and using Calibrated Clipping to match BF16 from 8B to 32B with up to 1.5x training throughput

Synopsis

This work studies the training stability of full-pipeline FP8 reinforcement learning for LLMs, identifies that compounded FP8 quantization noise distorts the importance ratio so that negative-advantage tokens are over-clipped at the lower bound and their gradients are erroneously zeroed, and proposes Calibrated Clipping, which dynamically realigns clipping bounds by matching the lower-bound clipping quantile of a BF16 reference and rebalancing the upper bound, eliminating entropy surges and restoring near-BF16 performance across GRPO and DAPO, 8B to 32B models, and multiple FP8 scaling granularities, while reaching up to 1.5x the BF16 training throughput.

AI-generated editorial illustration: Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Interpretation

The paper identifies a previously overlooked failure mode in full-pipeline FP8 RL: compounded quantization noise is amplified in the importance ratio, causing systematic over-clipping that triggers mid-training entropy surges and garbled outputs. Prior work mainly addressed the train-rollout precision mismatch with corrections such as TIS; this work locates the problem in the clipping mechanism of the PPO-style surrogate objective itself being distorted in quantized space, showing an independent source of instability that persists even when a unified FP8 pipeline narrows the precision gap. Based on controlled experiments training Qwen3-8B-Base on DeepScaleR with GRPO at 16K context, comparing a BF16 baseline against FP8 rollout paired with BF16, tensorwise, rowwise, and blockwise training backends; BF16 shadow passes measure the NMAE of probabilities and ratios, reporting that ratio noise is enlarged by 1.7x to 2.9x, alongside clipping fractions, entropy bins, and statistics on the origin of over-clipped tokens.

The paper proposes Calibrated Clipping: it first determines the lower clipping quantile from a BF16 reference and inverts the FP8 distribution to obtain a calibrated lower bound, then solves for an upper bound so that the positive-to-negative update contribution ratio matches BF16, with periodic recalibration to amortize overhead. Unlike simply relaxing the lower bound or targeting a fixed positive-contribution threshold, this method derives the lower bound and target ratio from a high-precision reference and solves for the upper bound accordingly, restoring the intended trust-region semantics in quantized space. The method is given as an algorithm with concrete search intervals, step size, smoothing update magnitudes, and updates every 20 training steps; the paper reports that the calibrated lower bound does not fluctuate drastically once training is stable, allowing the extra BF16 forward passes to be amortized, and provides a sensitivity study on the update interval and bound initialization in the appendix.

Across GRPO and DAPO, model scales from 8B to 32B, and multiple FP8 scaling granularities, Calibrated Clipping eliminates entropy surges and restores performance to near-BF16 levels. Compared with uncalibrated FP8 training, the method lifts clearly degraded average scores back to levels comparable with the BF16 baseline, with blockwise FP8 reaching 58.6 versus 57.6 on Qwen3-8B-Base and 53.4 versus 51.9 on Qwen2.5-32B. GRPO experiments evaluate on eight reasoning benchmarks every 25 steps over 500 training steps and report the best average score; DAPO experiments report Avg@32 on AIME24 over 200 training steps; additional results are given on the Eurus coding dataset with TACO, APPS, and Codeforces, together with training metric curves and calibrated bound trajectories.

Full-pipeline FP8 yields throughput gains in the training phase: tensorwise scaling reaches up to 1.5x the BF16 training throughput, while blockwise scaling still provides roughly 10% to 20% improvement. The paper extends the efficiency argument from the FP8 rollout generation stage, which prior work mainly emphasized, to the training stage itself, benchmarking across 8B, 14B, and 32B models at 4K, 8K, and 16K sequence lengths. Training-phase throughput is measured with an offline TorchAO benchmark, which the paper explicitly notes excludes periodic BF16 reference passes; the paper also cites prior work reporting roughly 30% generation speedup from FP8 rollout to characterize the overall full-pipeline efficiency space.

Perspective

The result targets engineering and research settings that use PPO-style surrogate objectives (GRPO, DAPO) for RL post-training and aim to move both rollout and training to FP8; it applies to models from 8B to 32B and sequence lengths from 4K to 16K, and requires access to BF16 master weights and periodic forward reference passes. For algorithms using sequence-level clipping or batch-level trust-region designs, the paper explicitly lists extending calibration to the sequence level as future work.

Calibration depends on BF16 reference passes, and how that overhead changes at larger models or with more frequent recalibration is only addressed by an offline throughput benchmark that the paper says excludes the reference passes; the paper also notes its hyperparameters were not explicitly tuned, and the sensitivity study shows some score variation across update intervals and initializations. In addition, the failure of CISPO and BAPO in this FP8 setting suggests it remains an open question whether other clipping designs also need calibration, and calibration for sequence-level clipping has not yet been verified.

Sources