Skip to main content
Back to timeline
arXivSource publication:

Task-Progress Distillation Lets a 1.7B Agent Reach 72.4% Unseen Success on ALFWorld from 404 Demonstrations

Synopsis

The work introduces Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short task-stage label so the student selects actions by jointly scoring admissible stage–action pairs; on ALFWorld, a 1.7B student trained with 404 demonstrations reaches 72.4% mean unseen task success with either TPD or action-only supervision, versus 48.3% for a reasoning-trained student using constrained action selection, and explicit stages raise success from 48.0% to 67.7% at 200 demonstrations, with action-only catching up to 76.9% at 808.

Source-provided article image: Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
Figure 1 ·

Figure 1 : Offline training and closed-loop execution. (a) A frozen annotator reads each teacher-history prefix and pairs the current task stage with the demonstrated action for student SFT. (b) The student scores 3 ​ | 𝒜 ⁡ ( h t ) | 3|\mathcal{A}(h_{t})| stage–action candidates, and the harness executes the action from the highest-scoring pair. The next input includes the action and environment feedback; the predicted stage is logged only for analysis.

arXiv

Interpretation

Introduces Task-Progress Distillation (TPD): each demonstrated action is paired with a short label describing the current task stage, and the student selects actions by jointly scoring admissible stage–action pairs that a deterministic harness executes in the environment. Compared with imitating actions alone or imitating long reasoning traces, TPD replaces the long reasoning target with a short stage label and uses the stage as scoring context; labels come from a frozen annotator that uses only feedback available before each action. Full-parameter SFT of Qwen3-1.7B on six ALFWorld task families, three training seeds, four demonstration budgets (98/200/404/808 trajectories), with main evaluation over all 134 valid_unseen tasks and a 50-step limit.

Compact supervision alone trains effective small agents: action-only (C) and TPD (B) both reach 72.4% mean unseen success at 404 demonstrations and 76.9% at 808. Prior reported agents learned from demonstrations with larger models or more demonstrations; here a 1.7B model reaches this level from a few hundred LLM-teacher demonstrations and acts without teacher calls at evaluation. Means over three seeds; the teacher reference reaches 88.1% (118/134), leaving an 11.2-point gap at 808 demonstrations; comparisons with external methods are comparisons of reported systems under differing training and evaluation protocols.

Explicit task stages add benefit at an intermediate demonstration budget: at 200 demonstrations TPD reaches 67.7% versus 48.0% for action-only, a 19.7-point difference, with all three seeds favoring TPD by 29.9, 19.4, and 9.7 points. The stage benefit is not uniform across budgets: it is 5.5 points at 98 demonstrations and zero in the mean at 404 and 808, so its value is concentrated where demonstrations are limited. Paired outcomes show B alone succeeds in 89 of 402 seed–task evaluations at 200 demonstrations versus ten for C alone; at 808 each alone succeeds in 25 cases, so equal means do not imply identical success sets.

Shared-history analyses link part of TPD's local advantage to subgoal-transition decisions: on 38 frozen prefixes, fixing the matching stage yields a progress action in 83.3% of evaluations at 200 demonstrations versus 65.4% averaged over the two mismatched stages; 11 of 15 B-only progress cases are stage-sensitive, nine of them heating at the transform stage. The analysis compares next actions under identical interaction histories and adds constant-stage and shuffled-stage controls: shuffling labels lowers seen success from 55.2% to 37.6%, indicating that alignment with current progress matters. The shuffle control uses three seeds over 116 seen tasks; the shared-history analysis covers 72 prefixes, three budgets, and three seeds, and the authors describe these as exploratory, with contribution to episode success requiring closed-loop intervention.

Perspective

The results apply to settings with an admissible-action set, full interaction history, and tasks for which progress rules can be defined; the authors state that experiments cover one environment (ALFWorld) and one student model size (Qwen3-1.7B), with the annotator supplying task knowledge through hand-defined progress rules. For readers training small agents from a few hundred demonstrations, this offers a reproducible recipe: compact serialization, completion-only loss, candidate scoring, and a deterministic harness, plus checkpoint selection on a seen probe. The additional value of stage labels is concentrated at an intermediate demonstration budget, making it most relevant when demonstrations are costly to obtain and tasks require transitions between subgoals.

The overlap between stage sensitivity and B-only progress covers only three tasks and five prefixes, and only two of the eleven cases preserve B's selected action when switching from full-joint to global action-value-only scoring, so the overlap is sensitive to the scoring rule. The authors note that several factors change together in the comparisons (target length, fraction of action tokens, candidate count, label predictability) and that current results do not separate these contributions; the three seeds measure training variation on one fixed dataset per budget; the shuffle and shared-history analyses are exploratory studies on the seen distribution. Output-token counts describe trajectory text and exclude candidate-scoring computation, and end-to-end latency and total training cost remain to be measured. A reader working from a fast parse without tables and figures would need the original for the specific paired outcomes, seen diagnostics, and external comparison values.

Sources