Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Researchers show mathematically that omitting or truncating importance-weight correction in RLVR implicitly reweights the reward, and propose an SMC correction mechanism

This work shows mathematically that, in sparse-reward RLVR, omitting or truncating importance-weight correction for enriched rollouts implicitly reweights the defined reward, and decomposes the resulting gradient error into scale, rotation, and variance; the authors develop a sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over sample length into additive accumulation, and they apply this analysis to interpret a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, corrected and uncorrected, helps avoid collapse in sparse domains.