CIS recasts training-inference mismatch as a log-odds displacement and tops the five-benchmark average on all three MoE models
Synopsis
The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
Interpretation
The paper writes the mismatch between the training and inference engines as an additive displacement in log-odds and gives a ratio-displacement identity: the mismatch ratio equals a displacement factor times a confidence factor. Earlier methods threshold the importance ratio or quantities derived from it, but that ratio mixes two factors, the actual engine discrepancy and the token's own confidence; separating them makes the discrepancy itself measurable on its own. Based on large-scale measurements of the probabilities both engines assign to the same sampled token: on the dense control the displacement concentrates near zero, whereas on the MoE it shows a pronounced heavy tail that survives the mapping to the log-odds displacement.
Building on this characterization, CIS truncates positive displacements at a single constant threshold, which maps back to a ratio cap that tightens as token confidence increases. A fixed ratio cap such as TIS corresponds in displacement coordinates to a threshold that becomes increasingly permissive as confidence grows, so its truncation bias concentrates on low-confidence tokens; CIS treats a displacement of a given size the same way at every confidence level. Theory shows CIS replaces the unbounded second moment of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess; diagnostics report less than half the bias of TIS below a low-confidence threshold, a lower overall bias, and added variance still well below that of the exact ratio.
Across three MoE models (Qwen1.5-MoE-A2.7B, DeepSeek-V2-Lite, Qwen3-30B-A3B) and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average among the evaluated baselines. Compared under one training recipe against token-level (exact ratio, TIS, IcePop, KPop†), sequence-level (Seq-TIS, Seq-MIS, GSPO) and numerical (FP16) baselines, CIS has the highest mean in all three model blocks. All methods share the same training recipe and code path and differ only in how they handle the mismatch, with mean and standard deviation over three training seeds; the authors state that on several transfer benchmarks the margin over IcePop is smaller than the across-seed variation and do not interpret those individual columns as statistically separated.
Ablations support three design elements: upper-side-only truncation, truncation at the source of the mismatch, and a small numerical floor. Replacing the confidence-dependent CIS boundary with a fixed TIS cap lowers average accuracy, a two-sided band that also lifts small displacements drops below the base model, and removing the numerical floor also lowers average accuracy. Ablations vary one element of the operator at a time on Qwen1.5-MoE-A2.7B with the rest at default; the threshold sweep shows every positive threshold stays above the uncorrected run, with the default giving the highest accuracy.
Perspective
The result targets teams running decoupled RL post-training frameworks in which an inference engine samples rollouts and a training engine computes gradients, especially in MoE settings; the method is implemented as an elementwise operator that changes neither numerical precision nor routing records, so it can be combined with approaches that reduce the mismatch itself, such as routing replay or FP16 training. The theory gives upper bounds under bounded per-token gradients and independent token contributions, noting that for correlated contributions the sample size is replaced. The authors observe that any cap with a finite maximum admits a similar bound, so the theory explains why truncation is needed rather than ranking CIS against other bounded caps.
In the supplied text the numeric cells of the main results table and the ablation table are empty, so the specific per-benchmark scores and seed spreads cannot be checked here and only the relative ordering described in the prose can be used. The authors list the scope of their claims: only MoE models were trained and no dense model was trained; measurements and training use bf16 and specific engines, so other kernels or numerical formats can change the size of the mismatch and may require retuning the threshold; the threshold and floor were chosen from a preregistered grid on Qwen1.5-MoE-A2.7B and kept fixed for the other two models, while every baseline uses the default hyperparameters of its reference implementation; in asynchronous training the displacement also contains the policy lag of stale rollouts, which the analysis does not model separately; and the token-level correction fixes the conditional distribution of each token given its prefix, not the shift in the prefix distribution itself.
