SkillGym turns human skills into verifiable training environments, letting a 35B model reach 51.47% on SkillsBench and exceed reported scores of Claude Sonnet 4.6 and others
Synopsis
SkillGym introduces a framework that transforms human-written agent skills into executable, verifiable training environments: its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions; it constructs and releases 2,756 environments across 12 categories and collects 8,364 successful trajectories (averaging 49 tool calls and over 60k logged text tokens), supporting supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards; under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.
Figure 1: Performance on general-agent benchmarks.
arXivInterpretation
SkillGym turns human-written agent skills from external inference-time instructions into executable, verifiable training environments, with a skill-to-task pipeline that instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions. Previously skills were typically used as external inference-time instructions rather than internalized as reusable model capabilities; this work turns the skills themselves into training environments and verification signals. The abstract describes the pipeline's three stages (instantiating tasks, code-based checker verification, contrastive execution to assess skill dependence) and states these resources support supervised fine-tuning and reinforcement learning with outcome-based rewards.
The authors construct and release 2,756 environments across 12 categories and collect 8,364 successful trajectories from multiple models and harnesses, averaging 49 tool calls and over 60k logged text tokens each. This provides a large-scale, outcome-verified resource for skill training and evaluation, rather than a demonstration on a single model or task. The abstract gives explicit numbers for environments, categories, trajectories, and average tool calls and token scale.
Under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.1 with and without skills, respectively. These gains come from supervised fine-tuning on verified workflows, indicating skills can be internalized as model capabilities rather than only supplied at inference time. The abstract reports specific benchmark names and numeric improvements and states the evaluation is under Claude Code.
The 35B SkillGym-Agent reaches 51.47% on skill-assisted SkillsBench, exceeding reported scores for Claude Sonnet 4.6, GPT-5.4 Mini, and DeepSeek V4 Pro; without skills it also surpasses skill-assisted bases under Codex and Claude Code, suggesting reusable procedural competence. A 35B model exceeds larger or stronger reported baselines in the skill-assisted setting, and still outperforms skill-assisted bases when no skills are provided, pointing to internalized rather than merely prompted skills. The abstract gives the specific 51.47% figure and comparison targets, and states the no-skill performance relative to baselines, though the comparison scores are 'reported scores' rather than reproductions under identical experimental conditions.
Perspective
The work targets training and evaluation settings where LLM agents solve real-world tasks, and suits teams that want to consolidate human-written skills into model capabilities; its resources (2,756 environments, 8,364 trajectories) and training recipe target models and harnesses that support supervised fine-tuning and reinforcement learning with outcome-based rewards, with the abstract's evaluations run under harnesses such as Claude Code and Codex.
The abstract does not specify the task-type boundaries covered by the code-based checkers, how contrastive executions quantify skill dependence, or the sources of variation across harnesses; the 51.47% is compared against 'reported scores' rather than reproduced under identical conditions, so readers should still watch the comparability of these comparisons. In addition, the currently visible text is the abstract and page navigation, without the body's method details, ablations, and failure-case analyses, so a complete understanding of the training recipe and skill-dependence mechanism awaits the full text.
