AlignOPSD aligns long-horizon agent supervision to functional decisions rather than timestamps, beating GRPO by 5.5–8.7% on ALFWorld, WebShop, and Search-QA
Synopsis
The work identifies Decision–Timestamp Mismatch in on-policy self-distillation for long-horizon agents—privileged guidance may correspond to a decision at a different timestep, and a student decision may span multiple timesteps—and introduces AlignOPSD, which first rectifies privileged supervision using functional correspondences across sibling rollouts and then applies semi-Markov hierarchical credit assignment over variable-duration decision spans; with Qwen2.5-3B/7B it outperforms GRPO on all eight backbone–aggregate-metric comparisons across ALFWorld, WebShop, and Search-QA by 5.5–8.7%, ranking first in six.
Interpretation
The paper identifies and characterizes Decision–Timestamp Mismatch: timestamp-local privileged evidence need not correspond to the student's functional decision, and the same decision can occur at different positions across rollouts, misaligning both the supervisory context and the credit horizon. Prior agentic OPSD work advanced along two lines—enriching teacher evidence via hindsight, execution feedback, or privileged contexts, and refining how teacher–student discrepancies are filtered and aggregated—yet its local signals remain anchored to positions along the student trajectory, assuming that teacher evidence at the current position is an appropriate basis for evaluating the functional decision there. Diagnostics on 100 WebShop tasks with Qwen2.5-3B-Instruct, four student and four privileged-view rollouts per task under a 15-turn limit: among positions where both trajectories are active and actions parseable, action types differ at 58.90% of student–student and 55.06% of student–privileged positions; under strict entity–action matching, 79.72% (S–S) and 81.31% (S–P) of corresponding actions occur at different turns, with a median displacement of four turns in both groups.
AlignOPSD proposes a two-stage framework of aligning supervision before assigning credit: Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts to calibrate local teacher evidence, and Semi-Markov Hierarchical Credit Assignment derives variable-duration decision spans from correspondence changes and allocates outcome-grounded credit across spans and their constituent turns. Unlike StepOPSD's action-centered steps or GEAR's adaptive regions derived from diagonal privileged divergence, AlignOPSD first resolves cross-rollout functional correspondence, then defines decision spans from changes in that correspondence and allocates a conserved outcome advantage across spans and turns. Method details appear in the main text and appendices: a frozen Qwen3-Embedding-0.6B encoder, similarity threshold 0.80, at most 3 sources with at most one per sibling rollout, maximum interpolation 0.80, and fallback to the identity gap when no valid source exists; span lengths constrained to 2–8 turns, credit temperature 0.5, boundary quantile 0.80, and Jensen–Shannon divergence between adjacent profiles.
Across ALFWorld, WebShop, and Search-QA with Qwen2.5-3B and 7B backbones, AlignOPSD outperforms GRPO in all eight backbone–aggregate-metric comparisons by 5.5–8.7%, ranking first in six and second in two. Gains hold over both trajectory-level RLVR (GRPO) and step-level self-distillation (StepOPSD), indicating the improvement extends beyond step-level supervision; because skill-conditioned baselines share the same SkillBank and retrieval source, the results suggest privileged information alone is insufficient and that behavioral alignment and outcome-conditioned credit assignment also matter. Reported values: ALFWorld average success 80.5%/89.1%, Search-QA accuracy 45.1%/49.1%, WebShop Score 86.8%/87.9% (3B/7B); improvements over GRPO are 5.5%/7.9% on ALFWorld, 8.7%/7.1% on Search-QA, 7.0%/7.0% on WebShop Score, and 5.5%/6.3% on WebShop Acc. Only two comparisons are not first: 3B ALFWorld at 81.2% trails GRPO+OPSD, and 7B WebShop Acc at 81.2% trails Skill-GRPO∗.
Ablations and mechanistic analysis support the separate contributions of the two alignment stages: removing supervision rectification lowers WebShop 7B Score/Acc, and replacing adaptive spans with token-, turn-, or random-level allocation also underperforms; diagnostics show cross-rollout correspondence is selective, span lengths are task-dependent, and rectification materially changes token-level supervision. This provides stage-wise evidence for the principle of aligning supervision before assigning credit, beyond end-to-end metric gains. On WebShop 7B, removing rectification drops Score from 87.9% to 84.2% and Acc from 78.9% to 71.9%; relative to token-level allocation the method improves 6.3/7.8%, relative to turn-level 1.6/1.6%, and relative to random allocation 9.0/9.4%. Learned mean span lengths are 4.41 turns on ALFWorld, 3.92 on WebShop, and 2.24 on Search-QA; off-turn sources account for 96.70%, 83.63%, and 22.71% of matching weights during training.
Perspective
The results target multi-turn interactive agent tasks with verifiable terminal rewards and training-time privileged information, such as embodied household control (ALFWorld), search-augmented question answering (Search-QA), and online shopping (WebShop); privileged inputs are not needed at deployment, and teacher and student share the same policy parameters, differing only in conditioning context. For researchers and engineering teams seeking better dense-supervision localization, the paper offers a reusable two-stage pipeline (functional-correspondence rectification plus hierarchical credit assignment), released code, and empirical operating regions for retrieval top-k, thinking-similarity threshold, and credit temperature.
The sensitivity analysis reports validation success curves from development histories, and the authors state these characterize empirical operating regions rather than across-seed uncertainty; the comparison between learned span lengths and functional spans currently has only aggregate-resolution annotation, so the paper does not report boundary F1 or claim the learned spans exactly recover semantic goals. The validity of thinking similarity as a decision-state surrogate is audited in an appendix with a method-faithful comparison, but the mean similarity of retained pairs partly reflects the selection criterion itself. In addition, this reading is full text, but tables and figures appear as text, so some graphical detail of individual values cannot be fully reconstructed from the text.
