RELACE scores actions by comparing original-context and outcome-conditioned likelihoods, reaching the highest reported mean success among compared methods on ALFWorld and WebShop
Synopsis
RELACE is a critic-free framework that uses a frozen behavior policy to teacher-force score each executed action under both its original context and an outcome-augmented context, normalizes and clips the ratio of these length-normalized scores within each trajectory to reweight discounted returns, and then builds local advantages by comparing weighted returns among actions from matched states, combining them with trajectory-level GRPO supervision; on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct it reports the highest mean aggregate metrics among the compared methods, including GRPO, GiGPO, HCAPO, and GraphGPO.
(a) ALFWorld Success Rate
arXivInterpretation
It introduces an explicit original-versus-outcome-conditioned scoring mechanism: the same executed action is teacher-force scored by the frozen behavior policy under both contexts, and the ratio of length-normalized scores is normalized by its trajectory mean and clipped to yield a retrospective relevance weight. HCAPO's practical estimator normalizes each hindsight score by the mean hindsight score along the trajectory and does not explicitly evaluate each executed action's original-context denominator; RELACE puts the original-context score explicitly into the ratio, separating baseline plausibility from outcome sensitivity. The mechanism is given in full formulaic detail in the method section and is contrasted in a WebShop ablation against a hindsight-score-weighting variant and a focal-only ratio-weighting variant, with the reported success-rate ordering supporting explicit original-context correction.
The retrospective weights reweight discounted task returns before a state-conditioned group comparison, producing local advantages that are added to a trajectory-level GRPO advantage, yielding fine-grained credit without auxiliary value or reward models. State-conditioned methods such as GiGPO compare unweighted discounted returns whose targets still reflect subsequent actions and transitions; RELACE lets outcome information enter the return targets before local comparison, so focal targets and group baselines use the same weighted, smoothed quantity. The ablation shows the focal-only variant performing worse while full RELACE, which weights both focal and baseline consistently, performs best, supporting the design motivation; main experiments report aggregate metrics across two model sizes and two environments.
It adds temporal smoothing and a success-protecting asymmetric mask to stabilize the local signal: adjacent-step smoothing shares the weighted return signal with the preceding action, and the mask removes all negative local corrections on successful trajectories while leaving the macro advantage unchanged. These stabilization components follow HCAPO but are integrated on top of likelihood-ratio-weighted return targets, making the local signal more usable under sparse terminal rewards. The method section defines smoothing and masking explicitly and notes that singleton groups and constant-target groups yield zero local advantage; training diagnostics report the temporal standard deviation of the mean retrospective weight falling by about an order of magnitude from early to late windows while mean hindsight scores rise.
Evaluated on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct, it reports the highest mean aggregate metrics among the compared methods, with clearer early learning-curve separation from GiGPO at 1.5B. Improvements extend beyond trajectory-level GRPO to the listed step-level and hindsight-based baselines, indicating that retrospective likelihood correction can complement structured local comparisons. Table 1 combines RELACE results with published baseline configurations (GRPO, GiGPO, and GraphGPO following Cheng et al.; HCAPO following Tan et al.), and RELACE and GiGPO validation histories are compared separately at shared checkpoints; an overhead analysis reports hindsight scoring as a modest share of the full logged iteration time.
Perspective
The work targets partially observable, sparse-terminal-reward multi-turn language-agent training, and is meant for researchers and practitioners who want action-level credit without auxiliary value or reward models and without additional autoregressive rollouts. It is validated on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct, and it provides a stage-wise timing profile for WebShop at 1.5B on one node with eight H100 GPUs, which helps assess the added cost inside the training pipeline. State matching is approximated in practice with a string-similarity threshold, so applicability is tied to the textual form of the observable representation.
The likelihood scores are explicitly described by the authors as practical credit proxies rather than exact hindsight posteriors, so they represent retrospective relevance rather than causal contribution, and their relationship to action utility is left to future work. State matching under partial observability uses approximate string matching, and matched observations need not identify the same latent state, so the robustness of this step remains an open question. Many specific numbers (such as baseline success rates, ablation variant success rates, and training-diagnostic values) are absent from the provided text, so exact figures cannot be restated here; precise comparisons require the original tables and figures. In addition, the end-to-end training-time difference relative to GiGPO would require a matched run with hindsight scoring disabled.
