ExpVoyager lets a skill curator navigate raw experience on demand, lifting ALFWorld success from 41.4% to 63.7%
Synopsis
The work reframes agent skill synthesis as a dynamic navigation problem over past experience and proposes ExpVoyager, in which a skill curator uses a Navigable Interface to inspect raw trajectories across views and resolutions while a Navigation State tracks reusable procedural knowledge and open questions, producing task-specific skills for a frozen executor; across ALFWorld, WebShop, and ScienceWorld it improves task success and reduces execution steps over no-skill and existing skill-based baselines, with gains that grow as the experience space scales.
Interpretation
The paper shifts skill synthesis from abstract-first-then-use to on-demand navigation: the curator actively decides which part of past experience to inspect and at what resolution for the current task, rather than committing experience to fixed procedural knowledge before downstream demands are known. Existing approaches (AWM, ReasoningBank, Trace2Skill) abstract past experience into fixed skills or memories before future demands are known, and SkillTTA synthesizes at test time but relies on top-k similarity retrieval. ExpVoyager instead defines experience access itself as an active decision-making process. Two preliminary analyses motivate this: 80 execution-verified oracle skills on ALFWorld plus 300 sampled source trajectories with oracle knowledge annotations show that even the best existing method fails to preserve the majority of oracle knowledge, and that existing retrieval paradigms achieve low knowledge-level recall while larger k trades marginal recall gains for rapidly dropping precision.
ExpVoyager has two components: a Navigable Interface that organizes each trajectory into trajectory-level views (action sequence, outcome, task metadata) and step-level views (observation, recorded reasoning, action, immediate result) with search_exp and inspect_traj operations, and a Navigation State updated each round that links experience access to target-relevant knowledge and open questions. Instead of treating a whole trajectory as a single retrieval unit, the interface lets the curator inspect evidence at the resolution the current investigation needs and expand to full trajectory context when required; the Navigation State interprets new observations into transferable procedural relations, conditions still to verify at execution time, and questions to investigate next. Ablations show ALFWorld success drops from 63.7% to 55.3% without the Navigable Interface and to 58.9% without the Navigation State; analysis against oracle annotations indicates removing the interface lowers knowledge recall and raises irrelevant knowledge, while removing the state slows recall growth through repeated visits to duplicated knowledge.
Across three interactive benchmarks, ExpVoyager consistently improves downstream task performance and reduces execution steps, and remains effective across different curator-executor configurations. The paper reports ExpVoyager outperforming the no-skill ReAct base and four experience-reuse baselines on ALFWorld, WebShop, and ScienceWorld, and Table 2 shows a smaller curator improving a much larger executor. Table 1 reports concrete numbers: with a Qwen3.5-9B executor, ALFWorld success 63.7% (base 41.4%) at 25.9 steps, WebShop success 28.7% with score 61.2, ScienceWorld success 37.4% with score 47.0; with a Gemma4-31B executor, 69.6%, 37.9%, and 55.6% respectively. Table 2 shows ALFWorld success reaching 71.3% and 75.9% with other curators.
ExpVoyager keeps gaining as the experience space scales and can synergize with pre-constructed skills while using the experience-access budget more effectively. In the online setting, gains are limited and performance can even degrade when experience is scarce, but as accumulated experience broadens ExpVoyager keeps improving while baselines plateau and the gap widens; in a controlled offline setting, baselines decline as more source experience becomes available while ExpVoyager improves. Combined with pre-constructed skills, the variants mostly raise task performance while reducing task-time cost. The online experiment is averaged over three random task orderings; in the budget comparison, ExpVoyager benefits from additional budget up to about 20 rounds before declining beyond roughly 20-25 rounds, whereas iterative top-k retrieval extensions perform worse under comparable token budgets and gain little or even lose from more budget.
Perspective
The results target text-interactive agent tasks: ALFWorld household interaction, WebShop online shopping, and ScienceWorld scientific reasoning, with a frozen executor and a curator producing task-time guidance as skill.md. The default configuration uses Qwen3.5-9B as both curator and executor with a 20-round navigation budget and averages over three random runs; the offline setting uses fixed source pools (1000 ALFWorld trajectories, 1000 WebShop trajectories, 500 ScienceWorld task-variation pairs), while the online setting accumulates experience along a task stream. The paper also demonstrates combining with pre-constructed skills such as AWM, ReasoningBank, and Trace2Skill by using an existing skill as initial knowledge and letting ExpVoyager navigate raw experience for on-demand refinement, which offers an incremental upgrade path for harnesses that already maintain skill libraries.
The paper itself reports several open issues: in the online setting gains are limited and performance can even degrade when experience is scarce, with advantages appearing only after experience accumulates; performance declines beyond roughly 20-25 navigation rounds, suggesting continued navigation can introduce redundant or less relevant experience. Failure cases in the case study show that even when ordering constraints or preconditions are identified during navigation, the final skill may lose action ordering, fail to recheck conditions after execution state changes, or fail to translate a precondition into concrete executable actions. In addition, the preliminary analyses are conducted on ALFWorld with execution-verified and manually reviewed oracle skills, yet the paper notes successful outcomes can still arise from task-specific shortcuts, and knowledge preservation and retrieval are measured with model-assisted 1-5 semantic matching and normalization. These are scope and open questions rather than faults of the work.
