Skip to main content
Back to timeline
arXivSource publication:

ActFirst-OPD lets multi-turn agents act before reasoning, speeding on-policy distillation training 2.3x, 1.8x and 4.9x on ALFWorld, WebShop and ScienceWorld

Synopsis

The work proposes ActFirst-OPD, which decouples environment interaction from full-response generation: the student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, switches to autonomous next-action prediction once the transition deviates from the reference trajectory, and asynchronously generates full think-then-act responses from the collected interaction contexts for token-level teacher supervision; across 0.6B, 1.7B and 4B Qwen3 students it achieves average wall-clock training speedups of 2.3x on ALFWorld, 1.8x on WebShop and 4.9x on ScienceWorld over Vanilla OPD, while matching or exceeding mean task success rate in eight of nine benchmark-model settings.

AI-generated editorial illustration: Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Interpretation

The paper identifies and characterizes a reasoning-blocked environment-transition bottleneck in multi-turn on-policy distillation: standard think-then-act rollouts complete lengthy reasoning before each short action, delaying environment transitions and experience collection. Prior multi-turn OPD work (e.g., TCOD using temporal curricula against trajectory-level KL instability, TurnOPD combining adaptive rollout depth with turn-level loss weighting) mainly addresses supervision quality and stability, whereas this work poses the latency of online experience collection itself as a distinct problem. The authors' profiling of Qwen3-1.7B on ALFWorld shows reasoning accounts for 81-95% of per-turn rollout time, and controlled comparisons show direct-action rollouts exhibit more unproductive repetition and yield fewer successful trajectories per 100 environment transitions.

The paper introduces ActFirst-OPD, which generates fast actions via reference-conditioned inverse dynamics plus an autonomous next-action prediction fallback, and makes full-response generation asynchronous, removing reasoning from the critical path of environment transitions. The method adapts the RIDM-style conditioning scheme, inferring actions from the learner's current observation and the expert's next observation, to multi-turn OPD, and adds an autonomous next-action prediction fallback after divergence plus asynchronous full-response generation that does not use reference next observations. The paper derives an idealized rollout speedup and its upper bound (Proposition 1) and analyzes finite-concurrency constraints in the appendix; ablations show removing inverse dynamics lowers mean success rate by 17.31, 1.00 and 3.04 percentage points on ALFWorld, WebShop and ScienceWorld, and removing the fallback lowers it by 1.05, 3.11 and 2.82 points.

Across three benchmarks and three student sizes, ActFirst-OPD achieves average wall-clock training speedups over Vanilla OPD while matching or exceeding mean success rate in most settings. Compared with Vanilla OPD, TCOD-F2B and TurnOPD, the method matches or exceeds Vanilla OPD's mean success rate in eight of nine benchmark-model settings and attains the highest mean success rate among the compared OPD methods at all three student sizes on ALFWorld and ScienceWorld. On ALFWorld the gains over Vanilla OPD are 8.15, 9.85 and 2.67 percentage points for 0.6B, 1.7B and 4B; on WebShop, Qwen3-1.7B gains a training speedup with only a 1.00-percentage-point decrease in mean success rate; reported training wall-clock time excludes teacher training, reference construction and evaluation.

The paper further locates where the speedup comes from: both higher environment-transition throughput and fewer environment transitions, with reference-conditioned interaction improving rollout quality. The authors separate the throughput benefit of asynchronous generation from the task-progress benefit of inverse dynamics rather than reporting only end-to-end time. Under matched settings, ActFirst-OPD raises environment-transition throughput from 3.65 to 5.68 transitions per second and executes 30.45% fewer transitions over the full-training window (38,616 versus 55,525); fixed-checkpoint rollout-quality comparisons show it lowers the fraction of tasks with unproductive repetition, raises successful trajectory yield per 100 transitions and uses fewer rollout turns than direct action at every checkpoint.

Perspective

The results target multi-turn text-environment training settings with high-quality offline reference trajectories: the ALFWorld, WebShop and ScienceWorld benchmarks, Qwen3-0.6B/1.7B/4B students with a task-specialized Qwen3-8B teacher, and evaluation in which the student uses standard think-then-act interaction without reference information. The paper notes reference caches cover all training tasks and can be reused across student sizes, and that teacher training and reference construction costs fall outside the reported training wall-clock time, so they amortize over a fixed task set and environment. Realizing the speedup requires sufficient serving concurrency: the controlled experiment shows speedup rising as the concurrency cap goes from 16 to 256, while at low concurrency both response lengths yield speedups below one, indicating the gain depends on workload and hardware.

The paper itself notes that reference next observations currently serve only as rollout guidance and are disabled after divergence, and that future work could use reference transitions for auxiliary supervision or realign rollouts with their reference trajectories; settings where such references are difficult to obtain are beyond its scope. Appendix E notes that the interaction-context distribution induced by fast rollouts need not match that induced by the think-then-act policy at inference time, so the results support the design's practical value in the evaluated settings without establishing that exposure bias at inference time is eliminated or uniformly reduced. In addition, most cells of the main result tables (Tables 1 and 2) are empty in this evidence bundle, with concrete numbers coming mainly from the prose, the ablation table and the appendix, so readers wanting per-setting values should consult the original tables.

Sources