Skip to main content
Back to timeline
arXivSource publication:

SkillMorph evolves agent skills via trajectory-guided fault localization, raising pass3 on SWE-Skills-Bench from 67.96% to 88.37%

Synopsis

The work proposes SkillMorph, a skill-evolution approach based on trajectory-guided fault localization: it abstracts raw trajectories into state-transition sequences, aggregates failure and success coverage evidence across repeated runs, evolution loops, and tasks to identify suspicious actions, then traces those actions back to edit sites that may span multiple skills and artifacts before generating revisions; on SWE-Skills-Bench and CannBot the evolved skills achieve higher pass@3, trial-level accuracy, and pass3 than the original skills and four existing skill-evolution methods, and a deployment with an AI operator-development team produced six accepted skill-revision pull requests.

Source-provided article image: Trajectory-Guided Fault Localization for Agent Skill Evolution
Figure 1 ·

Figure 1 : Structure of the mcp-builder skill.

arXiv

Interpretation

SkillMorph explicitly determines the edit sites, i.e., the skill contents that should be revised, before generating any revision, and grounds this decision in behavioral evidence from execution trajectories. Existing skill-evolution approaches generally prompt an LLM to generate revisions directly from task outcomes, case-level scores, or whole trajectories, leaving the judgment of which behaviors are problematic implicit and unrecorded; ContractSkill achieves explicit localization but assumes a skill can be expressed as an ordered sequence of steps with postconditions and that a suitable verifier exists. SkillMorph instead localizes through actions, because natural-language guidance lacks the explicit execution semantics of code. The paper illustrates the localization chain with the mcp-builder skill: in Task 1 the agent imports a document through connection A but queries through a second in-memory connection B, and since an in-memory SQLite database is private to its connection, verification fails with no such table: documents (11/13 checks pass); in Task 2 the agent stops the server with SIGTERM and never checks shutdown on input close, so the check times out (16/17 checks pass). Both problematic actions follow a read of reference/node_mcp_server.md, so both are traced to that artifact.

Suspicious actions are identified by combining stability and generality: failure and success coverage are compared across repeated runs, adjusted by changes since the preceding loop, and then aggregated across tasks. Existing approaches typically attribute problems from limited evidence, for example SkillAdaptor attributes a failed trajectory to its earliest actionable step and other approaches prompt an LLM to reflect on a minibatch of rollouts, without a measure of how consistently a behavior is associated with failures. SkillMorph adapts the Ochiai form from spectrum-based fault localization and adds evolution-specific handling: coverage counts are adjusted by their change since the preceding loop, and adjustment precedes cross-task aggregation so that opposite trends across tasks do not cancel out. The paper gives formal definitions for task-level coverage counts, the historical adjustment weight, and the suspiciousness score, and specifies that an error introduced by an action and later recovered from in a successful run still counts as failure coverage, because the same error may remain unresolved in another run and cause failure.

On SWE-Skills-Bench and CannBot, skills evolved by SkillMorph achieve higher and more consistent success on held-out tasks and outperform four existing skill-evolution methods. Relative to the original skills, pass@3 rises from 82.45% to 95.51% (15.84% relative) and pass3 from 67.96% to 88.37% (30.03% relative) on SWE-Skills-Bench, while on CannBot pass@3 rises from 80.00% to 93.33% (16.67% relative) and pass3 from 46.67% to 93.33% (100% relative). Against the strongest baseline CAP-Evolve, SkillMorph solves 16 more tasks on SWE-Skills-Bench, with larger gains on consistently solved tasks (17 to 68 tasks, 3.47 to 13.88 points in pass3). Across the 490 SWE-Skills-Bench tasks, SkillMorph achieves 95.51% (468/490) pass@3 and 88.37% (433/490) pass3, with 92.24% (1356/1470) trial-level accuracy over 1,470 evaluation runs; on CannBot it succeeds on 14 of 15 tasks in every trial, yielding 93.33% on all three metrics. Fine-grained comparison shows 126 tasks improved and 9 regressed on SWE-Skills-Bench versus 122 improved and 36 regressed for CAP-Evolve, and among 333 originally consistently solved tasks SkillMorph retains 327 (98.20%) versus 309 (92.79%) for CAP-Evolve.

Ablations indicate that repeated execution contributes most clearly to task coverage, explicit edit-site localization to consistent success, and history the least; the method also saw real deployment in AI operator development. The paper decomposes skill evolution into testable components and quantifies the cost of removing each, and reports deployment outcomes from collaboration with a development team. The Single-run variant lowers pass@3 by 5.31 (26/490) and 13.33 (2/15) percentage points on SWE-Skills-Bench and CannBot respectively; removing localization lowers pass3 by 6.12 (30/490) and 33.33 (5/15) points; w/o History has the smallest effect. In deployment, evaluation covers 63 AI operators and all 6 submitted skill-revision pull requests were accepted; for tasks passing correctness checks under both original and evolved skills, average completion time falls from 123.63 to 72.81 minutes (41.11% reduction); across 1,788 case-level measurements, 44.18% of generated kernels match or exceed the reference implementation.

Perspective

The method targets code agents that use file-based skill packages and whose tasks carry executable verification or tests, in settings where a skill set serves a class of related tasks and held-out tasks exist for evaluation. The paper validates on 49 independent skills from SWE-Skills-Bench and 12 coordinated skills in CannBot, covering repository-level software engineering and AscendC kernel generation, and deploys on 63 AI operators in an AI operator-development team. For a reader, this means that if your agent workflow can provide repeated execution and pass/fail judgments, the pipeline of localizing suspicious actions, tracing edit sites, and generating and filtering revisions is reusable; if skills cannot be judged by a verifier, or tasks do not form one class, the evidence reported here does not directly cover that setting.

Several open questions remain for a careful reader. Suspicious-action identification relies on a deliberately narrow LLM judgment of whether an action introduced an error state rather than which action caused a failed outcome; the paper notes the latter remains difficult even for strong LLMs, so the stability of this judgment affects downstream localization. The historical adjustment and cross-task aggregation weights are given by formulas, but the paper does not report their sensitivity across different skill collections. The ablation shows w/o History has the smallest effect, indicating a limited marginal role for historical information in the current setting; whether it matters more over longer evolution horizons or with more frequent regressions remains to be seen. The paper also notes that experience accumulated through cross-task evolution may still leave task-specific difficulties unresolved: in bidirectional cross-validation, 21 tasks failed in all three evaluation runs under both the original skill and the skill evolved using the opposite task fold, and of these, five subsequently succeeded in at least one run when they themselves belonged to the evolution set, suggesting that how task-level adaptation should combine with cross-task experience is still an open direction. In addition, CannBot uses a single fixed split because of substantially higher execution costs, so the dependence of its conclusions on the particular split is not further examined.

Sources