Skip to main content
Back to timeline
arXivSource publication:

SkillSeek matches an LLM-mediated retrieval loop on 89 SkillsBench tasks with BM25 plus a small reranker, cutting per-trial spend from USD 51.30 to USD 27.54

Synopsis

The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.

AI-generated editorial illustration: SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Interpretation

Across a grid of the 192-skill curated pool and the 34K marketplace pool crossed with Qwen3.5-397B-A17B and MiniMax-M2.7 on the 89 SkillsBench tasks, deterministic retrieval reaches observed parity with the LLM-mediated loop: plain bm25 beats liu_refined on three of four settings (192/Qwen3.5 at 0.430 vs. 0.397; 192/MiniMax at 0.387 vs. 0.344; 34K/MiniMax at 0.346 vs. 0.329), trails only on 34K/Qwen3.5 (0.420 vs. 0.442), where Qwen3-Reranker-0.6B records the same 0.442. Prior literature outsources selection to the agent itself through an LLM-mediated loop that rewrites queries and refines candidates; this work supplies the deterministic-retrieval point of comparison, sharing a sparse-plus-dense first stage and diverging at the reranker. Each main-grid setting is a single 89-task trial; the authors re-ran one setting three times and measured roughly 1% standard deviation, so they treat differences below 2% as within noise and state the result as observed parity rather than a resolved advantage. liu_refined is a conservative reimplementation, and the authors note a fully faithful one would raise its pass rates.

The cost structure is decomposed explicitly: SkillSeek pays nothing on the retrieval side (bi-encoder and cross-encoder run on CPU at roughly 1.1 s median and 1.9 s at p95 per skill_lookup call), leaving agent-side spend at USD 27.54, within fifty cents of the USD 27.41 no-skill baseline; liu_hybrid inflates the agent loop by 44% (USD 39.43) and liu_refined totals USD 51.30. Separating accuracy parity from cost difference lets the choice of retrieval scheme be weighed by per-task token spend rather than by pass rate alone. Costs come from controlled runs under the same driver, backbones, and task budget, with a table splitting retrieval-side and main-loop components.

The authors propose a first-stage recall ceiling to explain why lightweight rerankers suffice on the curated pool while heavier methods only catch up at marketplace scale: on the 192 pool bm25 reaches R@5 = 0.546 with BGE-base matching it and the cross-encoder adding only 3.5%; on the 34K pool bm25 reaches 0.391 and BGE-base 0.379, with the reranker adding 5.4%, growing to 9.9% once Tool-REX v3 restructures the index. The pool-size split is attributed to a measurable mechanism quantity (remaining first-stage recall) and cross-checked against the same methods across pools and against the order reversal inside Liu et al.'s own two variants. The analysis compares first-stage retrievers by R@5, and the authors state that recall is not the quantity the conclusion is stated in, using the helpfulness-gap diagnostic to show the two can move in opposite directions.

Indexed text and candidate depth dominate the design space: the Tool-REX v3 four-tag profile beats name+description by 2.2%, appending the raw SKILL.md body drops a further 3.2%, stage-1 depth of 20 is best (halving costs 2.1%, doubling or quintupling costs 3.5% to 4.5%), and returning more than top-3 skills to the agent buys essentially nothing. Indexing content and retrieval depth are reported as reproducible ablations, including the concrete failure modes of two earlier tag schemes (freeform tags dropped pdf-excel-diff R@5 from 1.0 to 0.5 by omitting the literal excel/xlsx forms; library-name tags dropped court-form-filling to 0.0 at 34K scale). Ablations run on the 34K/Qwen3.5 setting, where the authors note higher trial parallelism depressed absolute pass rates, so relative changes are reported from a 0.360 reference.

Perspective

The result applies to deployments using the OpenHands harness, the 89 SkillsBench tasks with either the 192-skill curated pool or the 34K marketplace pool, and Qwen3.5-397B-A17B or MiniMax-M2.7 as the backbone; within that setting the authors position the deterministic IR recipe as a strong default to try first, with LLM-mediated retrieval a natural fallback where deterministic methods fall short. The retriever is exposed as an MCP server with three read-only tools (skill_lookup, skill_load, skill_list), so any MCP-supporting harness can adopt it without source-level changes, and indexed text plus stage-1 depth are reusable tuning entry points. The authors also point to a risk-aware extension that composes retrieval with a pre-load risk scorer to drop high-risk candidates before they reach the agent.

Each main-grid cell is a single 89-task trial, and the authors report roughly 1% trial noise and therefore treat differences below 2% as within noise, so per-setting rankings should not be read as settled; they state the headline result as observed parity and note that establishing equivalence in a stronger sense would require a pre-specified practical margin and repeated independent rollouts of the headline settings. liu_refined is a conservative reimplementation, and the authors note a fully faithful one would raise its pass rates, so the parity on 34K/Qwen3.5 should be read as the optimistic side for SkillSeek. On MiniMax-M2.7 roughly half the tasks per condition hit network or timeout failures, depressing absolute pass rates by 5% to 10%, so that backbone serves only as supporting evidence. The 192 pool was hand-curated so every task has a relevant skill, so its absolute numbers reflect both retrieval quality and curator selection. First-stage alternatives are compared by R@5 only, and recall and downstream pass rate can move in opposite directions. All conclusions are scoped to the tested tasks, harness, pools, and backbones, and transfer to terminal, web, and software-engineering agents and to independently built skill-retrieval benchmarks such as SkillRet remains open.

Sources