Skill2Env synthesizes 2,963 executable environments from skills, and 1.5K fine-tuning trajectories lift a 35B agent from 36.6 to 45.0 across seven benchmarks
Synopsis
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
Interpretation
Skill2Env explicitly decomposes agent capability demands into five dimensions—environment understanding, planning, skill usage, long-horizon consistency, and error recovery—and uses 100 reusable difficulty patterns to turn each dimension into concrete task challenges, with a task blueprint jointly constraining the task instruction, execution substrate, workspace, and rubric-based evaluator. Earlier skill-based environment synthesis work (such as SkillSynth, Terminal-World, FACET, and SKT) mainly expands task and environment coverage from skills or scenarios; Skill2Env instead makes the capability to be exercised the organizing axis, so the same skill can yield tasks with different capability orientations under different pattern combinations. The paper provides a five-dimension capability demand table and construction and hardening controls for 100 difficulty patterns, plus a case study turning a Docling document-processing skill into a quarterly vendor-reconciliation environment that maps blueprint challenges to difficulty patterns; the evidence is framework description and case material rather than a controlled experiment.
Iterative Task Hardening uses solver execution evidence to diagnose which challenges are already handled easily, then strengthens or adds difficulty patterns and revises the blueprint and environment in coordination, making tasks progressively harder. Unlike one-shot synthesis, this method treats solver performance as a difficulty adjustment signal and allows newly discovered generalizable challenges to be abstracted back into the pattern pool for reuse in later task design. On 500 paired task identities hardened in both rounds, the full-credit rate for DeepSeek-V4-Flash under the pi harness falls from 48.40% to 24.80% and then 15.40%, while mean assistant turns rise from 25.91 to 36.54 and 38.85; this is a paired comparison holding task identity fixed.
Supervised fine-tuning on 1.5K high-scoring trajectories generated from Skill2Env environments improves Qwen3.6-35B-A3B on all seven downstream benchmarks, raising the unweighted average from 36.6 to 45.0 (+8.4). The gains are not tuned to a single benchmark: SkillsBench +14.34, Terminal-Bench 2.1 +13.5, SWE-bench Multilingual +7.7, AutomationBench +7.34, VitaBench +5.87, Claw-Eval +5.51, and -Banking +4.47, indicating that supervision organized around general capability demands transfers across domains and interfaces. The main table covers seven benchmarks with closed-source frontier and open-weight model references and includes FACET as a skill-based environment synthesis baseline; fine-tuning uses only 1.5K trajectories for 5 epochs with a global batch size of 32.
Task collections from later hardening stages provide better training supervision under the same data budget and quality threshold: the seven-benchmark average rises from 40.58 to 42.51 and 42.87 across hardening rounds. This separates the question of whether tasks are harder from whether the supervision is more useful, indicating that hardening changes the quality of extractable training signal rather than only raising failure rates. Each iteration applies the same SFT quality filter and randomly samples 500 qualifying trajectories, holding training-set size and quality threshold fixed; the authors note that random rather than matched sampling was used because the set of qualifying trajectories changes across iterations.
Perspective
The framework targets agent post-training settings that require tool use and multi-step interaction, and suits teams that can obtain executable skills from public skill libraries and provision dependencies and runtimes in a Linux sandbox. It lets researchers design training environments along capability dimensions rather than only by domain or skill coverage, and feeds solver performance back into task design; for practitioners seeking to improve the general execution ability of small and mid-sized agents with a small number of high-quality trajectories, the paper's path is to synthesize environments from skills first and then filter high-scoring trajectories for supervised fine-tuning.
The paper reports results for a single backbone (Qwen3.6-35B-A3B) under fixed synthesis-model and solver configurations, so whether other backbone and synthesis-model combinations benefit similarly remains an open question; frontier closed-source models remain stronger on several benchmarks, so the capability ceiling and remaining headroom are still to be characterized. Hardening effects are validated on 500 paired tasks and 500 trajectories under a fixed budget, leaving larger-scale and longer-hardening behavior to be observed. In addition, the loaded text includes the abstract, main body, appendices, and a case study, but some tables and figures appear as text, and a few numeric details (such as the exact hardening-trigger threshold and learning rates) are not fully given in the text, so readers needing exact reproduction should consult the original.
