Skip to main content
Back to timeline
arXivSource publication:

Synthesizing Verifiable Skills from Code at Scale: A Column on Code2Skill and CodeSkillBank

Synopsis

The work introduces Code2Skill, a fully automated pipeline that abstracts source units from 19,769 actively maintained GitHub repositories into atomic-operation, composite-workflow, and recurring-pattern skill records, verifies each record through source-body-blind reconstruction and source-aware comparison, and builds CodeSkillBank with 1,006,822 accepted records, improving 57 of 72 protocol-matched evaluations by 11.7% on average and outperforming trajectory-derived skill banks on all seven shared benchmarks.

AI-generated editorial illustration: Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

Interpretation

It proposes a code-grounded skill synthesis paradigm and a runnable Code2Skill pipeline: source units are first scored on reusable intent, ordered steps, boundary conditions, transferability, interface sufficiency, and non-triviality, then extracted into atomic, composite, or recurring-pattern records that capture applicability, execution steps, invariants, failure cases, anti-goals, and supporting evidence. Compared with trajectory-based synthesis that requires interactions with specific environments and document-based synthesis that lacks executable evidence, this work uses executable, verifiable, maintainable code as the substrate for abstraction and verification, so skills can be obtained before an agent accumulates interaction experience. The paper provides a full problem formulation, a four-stage pipeline description, and appendix prompt templates and decision algorithms, and it is actually run over 19,769 repositories with more than 500 stars, yielding 1,006,822 accepted records.

It builds a scalable grounding check from source-body-blind reconstruction, source-aware comparison, and failed-case adjudication: a model regenerates the implementation from the skill record alone, a judge compares it with the original code, directly accepts consistent cases, and routes the rest to an adjudicator that separates unsupported skills from reconstruction failures. This offers a repository-level consistency filter that does not assume tests or formal specifications, and the authors explicitly frame it as a filter rather than a substitute for test-based verification. Human annotation reports that 92% of skill descriptions in the final bank are judged accurate, 80% of records are judged worth retaining, and 84% of directly accepted records support correct reconstruction; the rejection sample scores 32%, 28%, and no correct reconstructions.

It constructs and releases CodeSkillBank as a large-scale skill resource with retrieval-oriented feature tagging and purpose indexing, while retaining repository-, file-, symbol-, and source-span-level provenance. Compared with prior skill libraries, the resource emphasizes auditable, refreshable, and deprecable records, and separates a task-facing retrieval view from the evidence archive so skills can be maintained as source code evolves. Statistics show records are dominated by multi-step procedures (57.3%) and constraint reasoning (33.5%), with surface API calls at only 8.6%; after purpose indexing, atomic-only cards account for 65.35%, composite-only cards 31.12%, and mixed cards 3.53%.

It tests where skills should enter the workflow in both inference and training: planning-time guidance improves all eight shared evaluations, generation-time prompting helps DS4-Flash more consistently, and post-generation critique improves 57 of 72 evaluations with the eight-benchmark average rising for every model and reasoning mode; in coding RL, post-generation review raises resolve rate from 24% to 38% at the shared training step 150. Rather than reporting a single insertion point, the work treats when skills enter the workflow as an independent variable and shows skills are most consistently useful when guiding planning or reviewing a concrete candidate. The inference side covers 72 protocol-matched evaluations across nine model settings and eight benchmarks; the training side compares conditions from the same Qwen3-32B SWE-World checkpoint under a fixed reward and training procedure, and the authors note it reports a single checkpoint without repeated seeds or learning curves.

Perspective

The results target agentic systems that need procedural guidance on coding, terminal and system interaction, and mathematical and scientific reasoning, in settings where actively maintained repositories exist and skills can be injected at planning or review stages; the bank is built offline without downstream task trajectories, so it can be used before an agent accumulates interaction experience. For teams seeking to turn domain data into reusable, executable knowledge assets, the pipeline offers an end-to-end path from source selection and extraction to verification and retrieval organization, with records that can be refreshed or deprecated as source code evolves.

A careful reader would still watch how reconstruction-based consistency checking relates to test-based or formal verification; purpose indexing reduces redundancy but may displace a locally more relevant candidate, and its downstream effects vary across models and renderers; the RL experiment reports a single checkpoint at the shared training step 150 and does not address learning speed, convergence, or final policy performance; and the human-code and AI-code banks differ by only 0.50 percentage points overall while disagreeing on 16 tasks, so equivalence between the two sources is not established. In addition, this reading is based on the loaded full text and external story, so any later updates to figures or appendix details would need to be checked against the original.

Sources