Skip to main content
Back to timeline
arXivSource publication:

HarA reweights RLVR advantages via FGW barycenters, lifting average reasoning scores by up to 4.8% across three Qwen3 scales

Related research and updates

Synopsis

The work proposes HarA, which represents each sampled rollout as a distribution over token hidden states and positions, uses the Fused Gromov-Wasserstein barycenter of all rollouts sharing an outcome to capture the current policy's internal reasoning pattern, and measures the novelty of a reasoning element by its contribution to the FGW distance to that barycenter in order to reweight token-level advantages of group-based RLVR; an anchor-guided linearization turns the FGW problem into a Wasserstein form solvable by Sinkhorn, and experiments on Qwen3-1.7B/4B/8B with GRPO, Dr.GRPO and DAPO over MATH500, Minerva Math, AMC23, AIME24 and AIME25 report average gains up to 4.8% over GRPO while maintaining higher entropy and longer responses.

Source-provided article image: Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry
Figure 1 ·

Figure 1: An example comparing existing credit assignment methods and the proposed HarA .

arXiv

Interpretation

HarA represents the current policy's internal reasoning pattern with an FGW barycenter and defines reasoning novelty relative to it. Existing GRPO credit assignment relies on local signals inside a rollout, such as token position or entropy; HarA instead projects each rollout into a distribution over hidden states and positions and compares it with the FGW barycenter of all rollouts in the same outcome group, yielding global novelty relative to the current policy. The paper gives formal definitions of the FGW distance and FGW barycenter, and an interpretability study on MATH500 reports a sparse, block-diagonal mapping between barycenter supports and LLM-derived reasoning stages, with the FGW formulation attaining higher average likelihood scores than Wasserstein and Gromov-Wasserstein variants and improving monotonically as the number of sampled rollouts grows.

Anchor-guided linearization converts the expensive FGW problem into a Wasserstein problem efficiently solvable by Sinkhorn, with an error bound and complexity analysis. The original FGW formulation contains a quartic, non-convex GW term; the authors fix the second transport plan to pre-align the first and last tokens of each rollout as reasoning anchors, turning the problem into a convex Wasserstein form. The paper provides the linearization derivation in Lemma 3.1, an approximation-error upper bound in Theorem 3.2 governed by the magnitude of order-reversing entries in the transport plan, and Theorem 3.3 showing time and space complexity linear in rollout length; an ablation reports per-step training time dropping from 154 to 41.0 (seq) and from 175 to 43.6 (tok) with under 1% accuracy variation.

After reweighting token-level advantages by novelty, HarA consistently outperforms existing credit assignment methods across models, benchmarks and group-based RLVR methods. HarA is a plug-and-play framework applicable to GRPO, Dr.GRPO and DAPO, and supports both sequence-level and token-level granularity, with the token-level variant outperforming the sequence-level one on most datasets and models. Evaluation spans 3 model sizes, 5 mathematical reasoning benchmarks and 3 group-based RLVR methods, reporting average gains up to 4.8% over GRPO and up to 4.0% over credit assignment baselines; appendix results on Dr.GRPO and DAPO show consistent improvements.

HarA changes training dynamics: faster convergence, better final performance, and sustained higher entropy and longer responses. Prior methods are often reported to suffer entropy collapse; HarA preserves exploration by emphasizing novel reasoning behaviors. On Qwen3-8B-Base with MATH500, the paper tracks accuracy, response length and entropy, reporting faster convergence, better final performance, higher entropy than GRPO and longer responses; efficiency results report up to 13.6% additional total training time, up to 2.64x speed-up to the same accuracy, and an 8.7% accuracy improvement at equal training time.

Perspective

The result targets researchers and engineering teams doing LLM post-training with group-based RLVR, in settings where rollouts are grouped by verifiable outcome for the same prompt and multiple rollouts are sampled per group; training uses 12,000 MATH problems and evaluation centers on five mathematical reasoning benchmarks with Qwen3-1.7B/4B/8B. The method plugs into GRPO, Dr.GRPO and DAPO, supports sequence-level and token-level granularity, and comes with a linearization error bound and linear complexity, so it can be applied directly to standard RLVR training and can serve as a starting point for bringing optimal-transport measures into other advantage-reweighting schemes.

Reading the FGW barycenter as an internal reasoning pattern rests on the premise that hidden states encode abstract reasoning steps; the interpretability study on MATH500 relies on an external LLM to derive reasoning stages, so mapping quality depends on prompt design. The exact forms of the novelty measure and advantage reweighting, including the full expressions in Eq. (7) and Eq. (8), are affected by typesetting in the main text, so readers should confirm details against the appendix. Evaluation is concentrated on mathematical reasoning and the Qwen3 family, leaving behavior on other task families, other model families and longer-response settings to be observed; sensitivity to barycenter size, the FGW trade-off parameter and the entropy regularization coefficient is only partly explored in the appendix.

Sources