Skip to main content
Back to timeline
arXivSource publication:

SkillGym turns 184k community skills into verifiable training environments, and its 9B fine-tuned model beats a 397B untrained model on two benchmarks

Synopsis

SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.

AI-generated editorial illustration: SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Interpretation

SkillGym reverses the usual direction by starting from community-written skills and building environments around them, rather than deriving skills from an agent's own experience in fixed environments. Existing skill-use training approaches mostly derive skills from an agent's own experience in a small set of environments; SkillGym keeps skills external and trains the general capability to use public skills, spanning 3.5k skills across 18 domains, with learned behavior transferring to skills held out from training. The paper reports crawling 184,430 skill manifests from two sources, 51,131 fetchable after deduplication, 11,897 skills after selection and cross-source deduplication, 3,494 of which produced tasks, for 6,772 tasks total; selection requires no runtime network access, no GPU, non-destructive actions, and at most 300 files.

Task construction uses a two-stage builder-reviewer agent system to generate tasks across procedural execution, abductive diagnosis, constraint satisfaction, and partial-order planning, each with a reference solution and an executable verifier. The paper notes prior work centers on a single reasoning structure and that none combines several structures to synthesize verifiable, skill-grounded tasks; SkillGym pairs a validity gate (the untouched workspace must fail the verifier and the reference solution must pass it) with a quality gate that audits for information leakage and for checks too strict or too weak. The paper reports 5,191 passed and 1,581 unresolved tasks out of 6,772; median instruction length is 336 words and median verifier code is 281 lines; ablations show that at the same token budget, training on review-approved tasks beats validated-only tasks by 4.7 points on the test set, 6.7 on SkillEval, and 9.3 on Skill-Use-Bench completion.

Supervised finetuning on the collected successful trajectories improves skill-use performance across model families and sizes, and the gains extend to skills held out from training. The paper reports improvements in 22 of 24 comparisons, with average gains of 13.8 points on the SkillGym test set, 9.7 on SkillEval, 9.7 on SkillsBench, and 41.2 in Skill-Use-Bench SU; the 9B model outperforms Qwen3.5-397B-A17B on the SkillGym test set and SkillEval. Six backbones span three families and 2B to 122B parameters; by skill split, held-in success rises from 40.0% to 57.0% and held-out from 42.5% to 62.0%; with skills available the SFT model rises from 43.0% to 59.5% while the base model rises from 33.8% to 41.3%.

The largest behavioral change is that trained agents consult the skills they are given, and the gains hold across reasoning structures. The paper breaks the Skill-Use-Bench score into three components: Trigger rises from 28.0 to 96.0, Compliance from 32.5 to 49.0, and Boundary from 57.7 to 64.0; annotating three benchmarks with a shared rubric shows SFT improves every structure by 16.9 to 22.1 points on SkillGym and 15.4 to 16.4 on SkillEval, with no structure-specific effect after controlling for task difficulty. Gains also reach structures rare in training: rule application is the primary structure of only 12% of training tasks, yet SkillEval's rule-application tasks improve by 13.9 points; on SkillsBench, tasks whose skills provide reference documentation gain 28.0 points, while those requiring a domain method, a shipped tool, or modifying an existing system do not improve.

Perspective

The pipeline targets researchers and engineering teams who want to train skill-use agents, and applies where skills can run reproducibly offline without GPUs or network access and without destructive actions; the paper keeps skills external and trains a general skill-use capability rather than compiling a specific skill into weights. Task construction covers procedural execution, abductive diagnosis, constraint satisfaction, and partial-order planning, with outcome-based verifiers that accept alternative valid solutions. The paper estimates the cost at $5.7k to $8.4k, or about $1 per released task, excluding environment building, annotation, and evaluation, which gives a reference point for reproducing or extending the pipeline at a similar budget.

The paper names two open directions: gains on SkillsBench concentrate on skills that supply reference documentation, while skills built around a domain method or a bundled tool improve little; and training uses only supervised finetuning, with reinforcement learning on SkillGym environments, whose verifiers already provide outcome rewards, as a natural next step. The ablations use token-matched subsets much smaller than the full data, and all ablation models fall below the full SkillGym model, so the paper suggests reading them as comparisons within each pair. In addition, task structure annotation was done with GPT-5.4, and the paper reports that its annotated number of checks matches SkillEval's evaluator specification for 99 of 100 tasks.

Sources