Skip to main content
Back to timeline
arXivSource publication:

CARM replaces geometric-mean masking with absolute-value averaging, raising mean@16 by up to 3.13 points on AIME 2024/2025/2026 and BeyondAIME

Related research and updates

Synopsis

The work proposes Cancellation-Aware Response Masking (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging so that opposing probability changes cannot cancel, and proves that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries; in experiments on mathematical reasoning and code generation, CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking and increases average pass@1 across four code benchmarks by 2.88 points over the strongest evaluated baseline.

Source-provided article image: CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
Figure 1 ·

Figure 1: A real unfiltered rollout that GeoMean accepts and CARM rejects. This 8,619-token response has scores s GeoMean = 1.0025 ≤ 1.005 s_{\mathrm{GeoMean}}=1.0025\leq 1.005 and s CARM = 1.0248 > 1.02 s_{\mathrm{CARM}}=1.0248>1.02 . Top: positive and negative log-ratio mass and prefix deviations. Bottom: a 103-token window; red tokens have log ⁡ r t > 0 \log r_{t}>0 , green tokens have log ⁡ r t < 0 \log r_{t}<0 , and color intensity encodes magnitude.

arXiv

Interpretation

CARM is a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling within a response. The common rule uses the length-normalized geometric mean of sampled token probability ratios, whose signed log-ratios can cancel across positions and conceal substantial bidirectional policy drift; CARM replaces that rule with absolute-value averaging. The abstract states the method definition and motivation and describes the mask as deciding whether an entire response contributes to optimization; no implementation details or ablations are provided.

The authors prove that accepted responses satisfy a joint bound constraining both the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond the band boundaries. This gives a theoretical characterization of the mask's acceptance condition rather than only an empirical threshold choice. The abstract states that the proof exists but does not give the theorem form, constants, or proof technique.

On mathematical reasoning, CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking. The comparison targets the geometric-mean masking rule that CARM replaces, quantifying the gain from cancellation awareness. The abstract reports the improvement magnitude and the benchmark set but not per-benchmark numbers, model scale, or training configuration.

On code generation, CARM increases average pass@1 across four code benchmarks by 2.88 points over the strongest evaluated baseline. The gain is reported on a second task family beyond mathematical reasoning, supporting cross-task applicability of the mask. The abstract gives the average improvement and how the baseline was chosen, without naming the four benchmarks or listing per-benchmark results.

Perspective

The work targets LLM systems that use reinforcement learning for post-training, in settings where rollout and training engines differ and policy updates make sampled responses off-policy. Its direct users are researchers and engineers building and tuning LLM reinforcement-learning training pipelines, who can substitute CARM for the geometric-mean rule at the sequence-level masking step. The evidence described in the abstract covers two task families, mathematical reasoning (mean@16 on AIME 2024/2025/2026 and BeyondAIME) and code generation (average pass@1 across four benchmarks), so the scope of the conclusions should be read as post-training settings on these two verifiable task types.

The abstract does not state the theorem's concrete form or how the band parameter is chosen, nor does it give per-benchmark scores, model scale, training steps, or computational cost, so the stability of the gains across models and training budgets remains to be seen. Whether baselines beyond geometric-mean masking were compared and how absolute-value averaging affects training stability are also not elaborated. In addition, this reading covers only the abstract; figures, appendices, and implementation details in the full text are not included, so judgments about engineering reproducibility should rest on the paper's method section and experimental tables.

Sources