PivotOPD concentrates distillation on pivotal mistakes and the turns after them, lifting Qwen3-1.7B by 5.5 points over the strongest baseline on ALFWorld
Synopsis
Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
Interpretation
Failures often hinge on a single early pivotal mistake, and that mistake usually remains recoverable. Prior multi-turn OPD methods addressed error accumulation by reweighting turns or restricting which turns receive distillation, but left open whether failures hinge on one decisive mistake and whether the student can still recover from it; this work uses ALFWorld's symbolic oracle to recompute the remaining optimal trajectory length at every turn, defines a pivotal mistake as an action that lengthens it or makes the task unsolvable, and backs the claim with counterfactual replays. Among failed trajectories of Qwen3-8B, Qwen3-30B-A3B, and Qwen3-235B-A22B, 62%, 57%, and 56% contain a pivotal turn, 58% across the three models; the first pivotal turn typically arrives early, at a median of turn 2–4 out of 30, after which the models waste an average of 6–9 more turns; in counterfactual replays of Qwen3-8B, correcting the first pivotal turn raises replayed success from 22% to 74%, and leaving the mistake in place while forcing the oracle action at the next two turns still reaches 71% (about four points more with three forced turns), with replays reproducing recorded outcomes so the counterfactuals are exact.
Standard on-policy distillation does not repair pivotal turns, because the student almost never samples the recovery action. The work turns 'OPD fails at pivotal turns' from intuition into a measurement: it splits failures into those after a pivotal turn and those without one, and directly measures the frozen policy's probability of the oracle action at each pivotal turn. After training Qwen3-8B with standard OPD, the overall failure rate on held-out tasks falls from 78% to 62%, but failures after a pivotal turn fall only from 74% to 69%, with most of the gain coming from failures without a pivotal turn (4% to 2%); distillation moves the median probability of the committed mistake from 0.90 to below 0.01 and raises the median probability of the oracle action by ten orders of magnitude, yet it stays below 0.01 at every evaluated pivotal turn, so a group of eight rollouts is not expected to sample it even once.
Concentrating dense supervision at pivotal turns and the turns after them, and explicitly training recovery, beats prevention alone. PivotOPD has a teacher model read each trajectory in hindsight, select candidate turns, and name a gold action, treating a candidate as pivotal when the student's committed action disagrees; preventive distillation uses reverse KL on the recorded response to steer away from the committed mistake, while recovery distillation uses forward KL on self-teacher responses to raise the probability of recovery actions the student rarely samples, both entering a single PPO update alongside group-based RL. On ALFWorld, at least one teacher-detected pivotal turn falls within one turn of the oracle-labeled one in 62% of failed trajectories on average, at least twice the random baseline; in component ablations the full method averages 73.7% success versus 72.5% for preventive only, 64.7% for recovery only, 62.6% for hints injected at random turns, and 71.9% when a generic reflection prompt replaces the gold action (dropping 9.3 points on Pick2); the recovery-budget ablation shows the selected budget raises the best validation score by 1.2 points on ALFWorld, 1.2 on WebShop, and 0.9 on Search-based QA.
The gains reproduce across benchmarks, two student sizes, another model family, and software engineering. Against 13 baselines (including GRPO, OPSD, RLSD, SDAR, TurnOPD, TCOD, SOD, StepOPSD, AgentOPSD, Skill-GRPO, OPID, Skill-SD, and PivotRL), PivotOPD ranks first on all eight per-benchmark averages and stays best in the self-distillation setting where the student serves as its own teacher, suggesting much of the gain comes from where and how the teacher intervenes rather than from teacher capacity. With the Qwen3-1.7B student it improves over the strongest baseline by 5.5 points on ALFWorld (73.7 vs 68.2) and 5.9 points on Search-based QA (44.5 vs 38.6); with the Qwen3-8B student it reaches 93.0 on ALFWorld, 47.4 on Search-based QA, and 88.2 score with 81.9 success rate on WebShop, leading by at least 1.4 points; on WebShop the 1.7B student beats RLSD by only 1.2 in score but by 14.1 points in success rate; under self-distillation it leads the strongest baseline by at least 1.0 point on all three benchmarks; a Nemotron-3.5-SFT student raises its SWE-Bench Verified resolve rate by 3.2% versus 0.4% for standard OPD, closing roughly a third of the gap to the teacher.
Perspective
The result targets multi-turn language agents in replayable environments: ALFWorld, WebShop, and Search-based QA all reproduce recorded observations exactly, so the recovery budget can reach 3, whereas on SWE-Bench Verified the audited turn is the final committed action and the recovery budget is 0, leaving only preventive distillation. It suits teams using group-based RL plus teacher distillation, especially where successful rollouts are scarce and group-relative advantages are identically zero. The diagnosis relies on ALFWorld's symbolic oracle; on other benchmarks pivotal turns are only estimated by the teacher model. At inference the student needs no teacher, hints, or recovery, so no test-time computation is added; during training the self-teacher forward pass and recovery rollouts add cost, and on ALFWorld the recovery-rollout time grows by more than ten times from budget 1 to 3.
The pivotal-mistake rate and recoverability were measured on ALFWorld with a symbolic oracle, while other benchmarks rely on teacher estimates, so how general this structure is in more open environments remains open. A wrong teacher-named action becomes a wrong distillation target: teacher detection agrees with the oracle within one turn in 62% of failed ALFWorld trajectories, and in the remaining cases supervision lands away from the first oracle-labeled pivotal turn; when the student serves as its own teacher the method is still best, but its ALFWorld success rate falls short of the stronger-teacher setting. The recovery budget and recovery weight interact and are selected per benchmark on validation, so applying PivotOPD to a new benchmark requires a sweep. Every agent sees a history window of at most five turns, and the turns wasted after a pivotal mistake are counted under that window, so part of the waste may come from the limited history. In addition, some numbers in the loaded text are swallowed by typesetting in the abstract and body, so a few figures and table captions cannot be verified directly from the text; conclusions touching those positions should be checked against the original figures and tables.
