Skip to main content
Back to timeline
arXivSource publication:

RASO retrieves from an external skill corpus and rewrites it across harnesses, beating retrieval-free skill optimization on four agent benchmarks and two models

Synopsis

The work proposes Retrieval-Augmented Skill Optimization (RASO), a framework that retrieves section-level knowledge from a public skill corpus and adapts it to the target task and harness through Cross-Harness Adaptation, where RASI builds an initial skill without any agent rollouts and RASU iteratively refines the skill by retrieving missing knowledge guided by execution feedback; across four benchmarks (OfficeQA, SpreadsheetBench, ALFWorld, WebShop) and two models (GPT-5.6-Luna, Qwen-3.5-9B), RASO consistently outperforms baselines without retrieval-augmented initialization and updating.

AI-generated editorial illustration: Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

Interpretation

RASO treats an external skill corpus as prior knowledge throughout skill optimization, rather than retrieving only at test time or using external knowledge only once at initialization. Existing skill optimization methods largely refine skills from the agent's own rollouts and leave the accumulated public skill corpus unexploited; RASO lets retrieval serve both initialization and update. The paper compares on four benchmarks and two models and reports ablations: relative to the RFSI+RFSU baseline, swapping in RASU alone raises OfficeQA from 40.70 to 47.56 and SpreadsheetBench from 51.67 to 61.07; swapping in RASI alone raises them to 45.93 and 58.45; combining both reaches 49.03 and 63.33.

Cross-Harness Adaptation is the operation that makes retrieved knowledge usable: it strips source-domain-specific instruction and re-expresses the underlying procedure in terms of the target harness's objects, commands, and units. Directly reusing retrieved skills can be ineffective or even harmful because source skills were written for other tasks and other tool assumptions; the paper models this mismatch as an explicit rewriting step. Ablations show that adding adaptation raises RASI from 41.86 to 45.74 on OfficeQA and from 41.43 to 49.17 on SpreadsheetBench; under the same RASI initialization, RASU rises from 46.70 to 49.03 and from 57.86 to 63.33. SkillRouter, which retrieves without adaptation, falls below the no-skill baseline on several benchmarks.

RASI produces a strong initial skill at zero rollouts, while RASU uses execution feedback to identify failure modes and retrieve the corresponding knowledge for iterative refinement. The construction of the initial skill is folded into the skill optimization problem instead of assuming an externally supplied starting skill; the update stage is driven jointly by textual gradients and retrieval queries. With GPT-5.6-Luna, RASI improves over retrieval-free initialization (RFSI) by +5.63 (OfficeQA), +4.77 (Spreadsheet), +3.24 (ALFWorld), and +1.17 (WebShop); with Qwen-3.5-9B the gains are +5.61, +2.74, +5.72, and +10.73. RASO improves over the strongest competing method by +3.49, +6.31, +1.49, +1.07 (GPT-5.6-Luna) and +4.85, +1.79, +7.47, +11.30 (Qwen-3.5-9B).

There is a useful range for retrieval and corpus scale: performance improves as retrieved sections grow from 1 to 5 and slightly degrades at 10, while retrieving from only 1% of the corpus already beats no retrieval. The paper gives an empirical trade-off between knowledge coverage and retrieval noise in retrieval-augmented skill optimization, and attributes gains to adapting diverse external knowledge rather than to particular source documents. The paper reports the highest performance at k=5 on both OfficeQA and SpreadsheetBench, with slight degradation at k=10, and therefore sets k=5 for all experiments; the corpus-size study shows performance generally continues to improve as the corpus grows, with the largest gains on SpreadsheetBench.

Perspective

The result applies to agent settings where a natural-language skill text is the decision variable and model weights stay frozen, and where task and harness descriptions are available along with execution feedback and a validation set. It makes a public skill corpus a reusable prior: for skill builders who want to cut rollout cost, RASI offers a zero-rollout path to an initial skill; for teams with an existing optimization loop, RASU can replace the update step under the same initialization. The evaluation spans document-grounded question answering, spreadsheet manipulation, a text-based embodied environment, and web shopping, and the paper explicitly blocklists benchmark-identifying skills from the corpus to avoid retrieving skills written for the evaluation benchmarks.

The text is presented as abstract, method, experiment tables, and appendix; details not expanded include the exact wording of the adaptation agent's prompts, an evaluation of retrieval-query generation quality, and the failure boundary of Cross-Harness Adaptation when source skills differ greatly from the target environment. Corpus blocklisting uses keyword matching plus repository-level exclusion, so benchmark-related skills not caught by keywords may remain; the paper handles this conservatively but does not quantify residual risk. Test sizes also differ substantially across benchmarks (172 OfficeQA questions, 280 SpreadsheetBench tasks, 134 ALFWorld games, 500 WebShop instructions), so gain magnitudes should not be compared directly across benchmarks.

Sources