X-Tree mines reusable skills from trajectories into a tree, lifting success by up to 4.5%, 5.8% and 4.1% on WebArena, ScienceWorld and WebShop at matched data and budget
Synopsis
The work introduces X-Tree: raw actions are first canonicalized into typed tokens such as type⟨date⟩ or click⟨button⟩, then adjacent pairs are recursively merged under an X-Score that weighs recurrence, span length and the fraction of occurrences in successful episodes, yielding a deterministic reusable experience tree mined from a trajectory pool with zero LLM calls; X-Tree is then integrated into three training settings — offline RL with each node as a training instance under a step-matching reward plus a depth-scaled completion bonus, online RLVR with an adaptive skill bonus that supplies signal while the verifier is silent and anneals as it becomes informative, and on-policy self-distillation where a rendered X-Tree replaces the LLM-written skill bank as the self-teacher's privileged c
Interpretation
X-Tree compresses a trajectory corpus into a deterministic, auditable hierarchy of skills, where each internal node captures how a frequent and success-bearing skill is composed from sub-skills. Earlier agents that exploit hierarchy keep LLM-written natural-language skills in context and retrieve them; the skills never enter the weights. X-Tree instead recovers the hierarchy from the data itself by counting-based merging, with no LLM calls. The paper reports the mining procedure and hyperparameters (256 skills from 7,974 WebArena trajectories, 80 from 1,673 ScienceWorld trajectories, 48 from 1,824 WebShop trajectories) and grounded node examples, such as a depth-2 WebArena node that fills two fields and submits, occurring 411 times and appearing both as an admin-panel date filter and as a map route query.
In offline RL, where only trajectories exist and no action can be executed, treating each X-Tree node as one training instance with a step-matching reward plus a depth-scaled completion bonus beats full-data SFT at matched budget. Standard SFT and RLVR treat the action stream as flat and weight every token equally; this work instead assigns credit at node granularity and scales the completion bonus with depth. On the 694 deterministic WebArena tasks, the full X-Tree recipe reaches 22.9 SR against 18.4 for the Go-Browse SFT-Full recipe; matching RL compute with two extra SFT epochs raises SFT-Full only to 18.8, so the gain is not from compute; ablations put whole trajectory at 19.9, random span at 19.5, random tree at 17.9 and binary outcome reward at 18.2, all below X-Tree.
In online RLVR, X-Tree supplies an adaptive skill bonus that separates rollouts when a group has no success and the outcome reward is uninformative, then fades as successes appear. Unlike a learned process-reward model, this shaping signal is mined from the trajectory pool itself and its weight is tied to the group win rate. On ScienceWorld it leads outcome-only GRPO at every scale by up to 5.8% SR and leads on the held-out G2 fold at every scale by up to 5.8% SR; on WebShop it improves success by up to 4.1% and graded score by up to 4.1%; varying the rollout group size over 2/4/8/16 shows the gain is largest when rollouts are few and outcome signal is sparse, and narrows as rollouts grow.
In on-policy self-distillation, rendering X-Tree as the self-teacher's privileged context substitutes for an LLM-written skill bank at zero LLM cost. OPSD normally depends on an LLM-written skill bank as privileged context; this work replaces it with a deterministically mined and fixed-template-rendered X-Tree, so teacher and student differ only in context and any gap is attributable to that context. Across ScienceWorld and WebShop, several generalization levels and three model scales, the rendered X-Tree is comparable to the LLM-written bank (gpt-oss-120b on ScienceWorld, GPT-o3 on WebShop) and ahead in most runs by up to several points on each environment; ablations show an empty privileged context returns to the outcome-only level and scrambling the X-Tree words costs points.
Perspective
The work targets multi-step agent training where a trajectory pool exists but environments or verifiers are scarce: the offline RL integration fits domains with trajectories only and no executable environment (the 694 deterministic WebArena tasks), the online RLVR integration fits domains with an environment and a verifier (ScienceWorld, WebShop), and the OPSD integration fits settings that already have a self-distillation trainer and need a replacement skill bank. It lets researchers use the same trajectories more fully at matched data and budget, and makes the skill bank reproducible from data alone without LLM calls; the paper also notes that X-Tree marks where a corpus is thin, indicating which trajectories are worth synthesizing.
X-Tree is mined once from a fixed pool and is not improved during training; the paper lists mining from the policy's own trajectories so structure and policy improve together as future work. Each of the three integrations is evaluated in one setting and they are never combined, so the combined effect remains open. The OPSD gain depends on how well the base model reads and follows the skill text, and the paper observes larger margins at 7B with only a few points at 1.5B and 3B. In addition, this load contains the paper body and appendices, but some table cells appear as placeholders in the text, so per-cell means and standard deviations should be checked against the original tables.
