UniSkill trains a skill proposer with contrastive action feedback, reaching 98.4% success on ALFWorld with stable joint training
Synopsis
UniSkill lets a shared policy act as both the actor that executes environment actions and the proposer that emits skillbank edits (Add, Update, No Edit), and trains the proposer with contrastive action feedback that measures how replacing the retrieved skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories, avoiding extra rollouts per proposal, with skill-edit support regularization preserving exploration; it reaches 98.4% success on ALFWorld and 84.7% success on WebShop while sustaining stable joint training over 250 steps.
Figure 1: ALFWorld training dynamics. (a) UniSkill continues improving later in training, whereas Evolving-RL collapses after strong early gains. (b) The learned skillbank grows rapidly early and continues evolving through occasional additions and updates (non-overlapping 10-step means).
arXivInterpretation
Introduces UniSkill, a shared-policy framework in which one policy both acts on retrieved skills and proposes structured skillbank edits (Add, Update, No Edit) from completed trajectories, with skill-edit support regularization to sustain exploration of edit operations. Unlike prior work that extracts skills with a frozen LLM or trains the proposer from task-outcome prediction accuracy and trajectory rewards, this places actor and proposer under one jointly learned policy and explicitly addresses proposal-level feedback suppressing edit operations. Section 3 gives the full objective and regularizer; Section 5 Q3 compares skill-edit distributions and validation curves with and without support regularization, showing collapse to a single operation without it and continued sampling of all three operations with higher later success with it.
Develops contrastive action feedback: before the policy update, the current actor and recorded behavior are held fixed while only the conditioning skill changes, and the change in the action log-likelihood gap between successful and failed reference trajectories serves as an actor-alignment signal. Compared with Evolving-RL, which evaluates candidate skills through additional skill-conditioned rollouts, this reuses already collected trajectories and avoids per-proposal rollouts, reducing skill-evaluation overhead. Section 3.3 defines the likelihood and alignment reward; Section 5 Q2 samples 100 eligible proposals at training steps 25 and 75 with 32 rollouts per condition, reporting positive Spearman correlations between alignment scores and success-rate gains, while 24.5% and 15.6% of positive-score proposals still show negative gains.
Achieves strong performance on ALFWorld and WebShop: 98.4% success on ALFWorld and a task score of 90.5 with 84.7% success on WebShop, with stable joint training over 250 steps and an overall upward validation-success trend. Against RL-only baselines, ALFWorld success exceeds GRPO by 20.8 pp and GiGPO by 7.6 pp; against skill-augmented RL, WebShop success exceeds SkillRL by 12.0 pp; against the closely related joint-training baseline Evolving-RL, ALFWorld success is 5.4 pp higher, and Evolving-RL shows performance collapse under extended training. Table 1 reports main results; ALFWorld main and out-of-distribution results are means and sample standard deviations over three independent training runs, with baselines drawn from prior work or reproduced under a common protocol.
Ablations show joint training and actor-alignment feedback each contribute: the variant without the alignment reward reaches 84.4% without retrieval at evaluation, versus 96.1% for full UniSkill, and proposer-only training reaches 34.4%. Because the no-reward variant follows the same joint-training procedure and skillbank update rules and differs only in whether the alignment reward enters the proposer objective, the comparison attributes the actor's no-retrieval gain to actor-alignment feedback as a training signal. Table 2 lists ALFWorld success rates with and without retrieval for four settings; all variants start from an empty skillbank, retain skill retrieval during training, and use identical critic checks and skillbank update rules.
Perspective
The work targets text-based multi-turn interactive tasks: household navigation and object manipulation in ALFWorld and simulated e-commerce search and purchase in WebShop, evaluated on the official valid_seen split for main results and valid_unseen for out-of-distribution, with no real-world physical actions or purchases. The method suits settings that maintain a skillbank and need to credit skill proposals during joint training, especially where per-proposal environment rollouts are costly; the shared policy remains effective with a 3B backbone, indicating the mechanism is not tied to one model scale. For a reader, the directly reusable pieces are the contrastive action feedback evaluation idea and the skill-edit support regularization for sustaining exploration, with an implementation link provided.
Alignment scores correlate positively with rollout gains in aggregate, but 24.5% and 15.6% of positive-score proposals show negative gains at training steps 25 and 75, so the score is an informative proxy rather than a per-proposal guarantee; case studies separate three sources of disagreement: an added requirement causing exclusion, an actor not consistently following an otherwise reasonable new skill, and a proposal itself omitting the placement-before-search ordering constraint. Skillbank admission requires both critic checks and the alignment threshold, a stricter criterion that excludes some positive-score proposals. Skillbank growth is rapid early and slows as accepted Add operations become less frequent, and long-horizon evolution is covered mainly within the 250-step training range.
