Skip to main content
Back to timeline
arXivSource publication:

GraphOPD reweights distillation supervision with an environment-state dependency graph, beating the strongest baseline by up to +5.8 pp on ALFWorld and WebShop

Synopsis

GraphOPD reads which steps enabled which later ones from the environment's own record of state changes, builds a directed dependency graph, scores each step by a random-walk stationary distribution, and fuses that structural credit with the teacher-student divergence into a trajectory-relative mask that distills only each rollout's above-average-aptitude steps; across ALFWorld, WebShop, and SearchQA at three model scales against eleven baselines it stays competitive throughout, improving over the strongest baseline by up to +5.8 pp.

Source-provided article image: GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
Figure 1 ·

Figure 1: Motivating empirical study. All runs distill only the k k highest-KL steps or k k randomly chosen steps per trajectory, k ∈ { 5 , 8 , 10 , 12 } k\in\{5,8,10,12\} , on ALFWorld with a 7B Qwen2.5-Instruct agent under GRPO for 150 steps, all other conditions identical. ( Left ) Random selection stays ahead at every budget k k , by + 1.5 +1.5 to + 3.2 +3.2 pp in success rate, with the per-budget gaps Δ \Delta annotated above each pair. ( Middle ) At k = 5 k=5 the KL-guided run trails in success rate (solid curves, left axis) throughout and ends at 76.3% against 83.5% for random, while its policy entropy (dashed curves, right axis) contracts as the random run’s expands. ( Right ) The per-step KL of 466 successful and 466 failed steps from 50 rollouts overlaps by 84% between the two outcomes, with medians (dashed) at 0.10 and 0.09, so success and failure are barely separable by divergence.

arXiv

Interpretation

The paper first tests the existing rule that allocates on-policy distillation supervision by teacher-student divergence magnitude, and finds it fails on multi-turn agent trajectories: distilling the highest-KL steps brings no consistent benefit over random selection, the divergence-guided rule trails random selection throughout training, and successful and failed steps barely separate by divergence. Prior work carried the single-turn intuition that a large disagreement marks a mistake worth correcting directly into multi-turn agents; this controlled comparison shows that premise does not hold over long horizons, reframing supervision allocation as a problem that needs a drift-immune signal. On ALFWorld with a 7B Qwen2.5-Instruct agent under the same GRPO backbone and training configuration, highest-KL and random selection are each trained for 150 steps; the outcome-level check uses 50 rollouts providing 466 successful and 466 failed steps.

GraphOPD reads which steps enabled which later ones from the environment's own record of state changes, builds a directed dependency graph with one node per step, scores each step by a damped random-walk stationary distribution with uniform teleport, fuses that structural credit with the divergence signal into a distillation aptitude, and standardizes it within each trajectory into a soft mask. The authors state this is the first method to bring graph-based structural augmentation into on-policy distillation for agents; prior graph or branching credit scores reweighted only the RL advantage and left the distillation loss unweighted, whereas this work represents for the distillation loss which steps the later ones actually depend on. Dependency edges come from a rule-based State Transition Footprint over action and observation text, weighted by entity rarity, with no LLM calls, privileged annotations, or rollout counterfactuals; the appendix proves existence, uniqueness, and a geometric power-iteration rate, and graph construction is measured at under 0.5% of per-step training time.

Across ALFWorld, WebShop, and SearchQA at three model scales against eleven baselines, GraphOPD is competitive throughout, attaining the best result on ALFWorld and WebShop at every scale, peaking at 91.9% ALFWorld success and 83.6% WebShop success at 7B, and improving over the strongest baseline by up to +5.8 pp. Against GEPO, which also uses a graph-based credit score but reweights the GRPO advantage rather than adding a distillation loss, the margin is +5.2 to +7.9 pp on ALFWorld and +3.6 to +5.5 pp on WebShop success rate; against SDAR it is +12.0 pp on ALFWorld and +13.5 pp on WebShop at 1.7B. GraphOPD, SDAR, OPID, and GEPO are each trained with three independent runs; GraphOPD's own maximum standard deviation is 1.60 pp and stays below 0.6 pp in seven of the nine headline cells, against the largest baseline spreads of about 5 pp; all remaining rows are single-run point estimates.

Ablations show the two signals are complementary and each independently necessary, with the structural credit signal the larger single contributor; an executed-replay audit shows the structural credit score reaches a Spearman correlation of 0.58 and Hit@10% of 0.48 with true causal impact, clearly above the random baseline and above the divergence signal; the same signal also transfers to out-of-domain tool-integrated reasoning. A Shuffled Centrality control keeps the mask's sparsity and weight distribution but randomizes which step receives which weight, and at 7B and on WebShop and ALFWorld at 3B it costs up to 1.6 pp more than Divergence-only, so the gain comes from the content of the structural credit rather than from non-uniform weighting itself. The audit samples 10 decision steps per trajectory for 30 trajectories of the trained 7B policy, re-samples 4 alternative actions at each, and rolls forward to the horizon; the out-of-domain test runs AIME24, AIME25, LiveCodeBench-v5, and GPQA-Diamond as tool-integrated reasoning with Qwen2.5-7B-Instruct, averaging 29.0% against SDAR's 20.0%.

Perspective

The result targets multi-turn agent post-training where interaction is text-interfaced and reward arrives only at the end of the trajectory, covering ALFWorld-style embodied manipulation, WebShop-style web purchasing, SearchQA-style retrieval QA, and Python-interpreter tool-integrated reasoning. The method uses GRPO as its backbone and the student's own frozen snapshot as teacher, acting as a drop-in reweighting that leaves the optimizer and rollout infrastructure unmodified and needs no extra model calls. Teams seeking to cut trial-and-error cost in step-level supervision allocation, or already using divergence-weighted distillation, can swap in this allocation rule; teams needing cross-domain transfer only need to supply a rule-based Footprint extractor for the new environment.

The structural credit score depends on rule-based Footprint extractors, and the authors list using an LLM to extract the Footprint directly from raw trajectories as a natural next step, so transferability to new action spaces depends on whether an extractor can cover that space's entities and state keys. The soft mask literally excludes no step, and its effective concentration depends on the shape of the within-trajectory aptitude distribution; the authors report that on two logged rollouts the mean weight of above-mean steps exceeds that of below-mean steps by about a factor of two and extreme steps differ by up to a factor of seven, and note a budget-matched hard-mask variant as a natural ablation left to future work. The audit covers 30 trajectories with 10 decision steps each, and out-of-domain transfer is verified only at 7B on four benchmarks. In addition, all rows of Table 1 other than GraphOPD, SDAR, OPID, and GEPO are single-run point estimates, so readers comparing those rows should note the missing variance information.

Sources