Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

HarA reweights RLVR advantages via FGW barycenters, lifting average reasoning scores by up to 4.8% across three Qwen3 scales

The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.