Public articles linked to the same research event.
arXiv The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.
The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.
The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.
The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.