Writing Agent Skills as Graphs: GraphSkillEvo Evolves Reusable Procedural Knowledge
Synopsis
The work represents LLM agent skills as graph-structured natural-language artifacts, where each node is an execution step with its operational guidance and directed edges encode context-dependent transitions between steps, and builds GraphSkillEvo, a population-based evolutionary framework with global-guidance mutation, graph-structure mutation, global-guidance crossover, and graph-structure crossover that searches this structured skill space, outperforming the skill-optimization baseline SkillOpt on average across five agent benchmarks, two LLMs, and two execution settings (+4.01% on GPT-5.4-nano and +1.76% on GPT-5.4) while consuming fewer total optimization tokens.
Interpretation
It formulates agent skills as graph-structured natural-language artifacts: each node is an execution step with self-contained operational guidance, and directed edges are defined by consecutive nodes in workflows, representing transitions between steps under a given applicability condition. Prior skill-optimization methods such as SkillOpt represent skills as unstructured natural-language instructions, which the authors describe as lacking explicit workflow-level guidance, carrying substantial redundancy, and leaving a vast search space; the graph structure organizes execution steps and their dependencies explicitly so shared instructions can be specified once and reused across workflows. The paper gives a formal definition (global guidance, node set, edge set, workflows) and states three properties: low redundancy, explicit workflow guidance, and a more compact structured search space; a validator script checks that generated skills follow the graph schema, and skills failing validation are discarded and regenerated.
It introduces GraphSkillEvo, a population-based evolutionary computation framework that searches the structured skill space with four structure-aware operators (global-guidance mutation, graph-structure mutation, global-guidance crossover, graph-structure crossover) and performs population selection by validation fitness. Compared with purely LLM-based iterative self-refinement (SkillOpt's patch-style updates with validation gating), maintaining a population and recombining components lets effective parts discovered along different search trajectories be combined, broadening exploration. The paper details the full procedure: population initialization, small-batch execution on the training set with failed trajectories retained as reflection information, round-robin operator selection, fitness-rank-based parent selection, evaluation of new skills on the full validation set, and retention of the highest-fitness individuals across generations before returning the best skill.
Across five benchmarks (SearchQA, SpreadsheetBench, DocVQA, LiveMathematicianBench, ALFWorld), two LLMs (GPT-5.4, GPT-5.4-nano), and two execution settings (no harness, Codex harness), GraphSkillEvo achieves the best result in 13 of 14 settings and beats SkillOpt on average. Relative to no-skill execution, average success rate improves by 15.37% (GPT-5.4, no harness), 21.86% (GPT-5.4-nano, no harness), and 10.31% (GPT-5.4, Codex harness); relative to SkillOpt the gains are 1.76%, 4.01%, and 1.33% respectively. Results are test-set success rates averaged over three repeated skill-optimization runs; the authors report the largest gains on procedural benchmarks (SpreadsheetBench +10.60% and ALFWorld +3.73% over SkillOpt) and provide one-sided Welch t-tests, with Spreadsheet and ALFWorld below p = 0.05 and SearchQA and DocVQA in the 0.05 to 0.10 range.
Ablations and transfer analysis support the value of the graph structure itself: removing the graph structure, mutation, or crossover all reduce average performance, and skills optimized with GPT-5.4-nano remain better than the no-skill baseline when deployed on GPT-5.4. This separates the contribution of the representation from that of the search mechanism: without graph structure the average drops from 71.52 to 64.08, indicating population-based evolution alone is insufficient; without mutation it falls to 54.50 and without crossover to 66.59. Ablations are averaged over three repeated experiments on GPT-5.4-nano; an execution-side control converts graph-structured skills into unstructured counterparts retaining global guidance and node-level instructions, with drops of 4.52, 2.50, 4.19, 1.35, and 0.75 percentage points across the five benchmarks; in transfer, the transferred GraphSkillEvo skill reaches 71.78 on SpreadsheetBench, above its directly optimized counterpart (69.40) and the transferred SkillOpt skill (53.21).
Perspective
The results target settings where natural-language skills improve LLM agent task success, covering five benchmark families (fact-based question answering, spreadsheet manipulation, visual document understanding, mathematical multiple-choice reasoning, and embodied interaction) under two execution settings (no harness and Codex harness); optimization changes only the skill text while model parameters and the execution harness stay fixed. For readers interested in reusing procedural knowledge, the value lies in cross-model transfer: skills optimized with GPT-5.4-nano still beat the no-skill baseline on GPT-5.4. The authors list future directions including combining the method with parametric optimization, extending to richer graph composition mechanisms, and merging graph-structured skills from diverse domains.
Worth noting is that gains are uneven across benchmarks: differences on procedural benchmarks (SpreadsheetBench, ALFWorld) reach statistical significance, question-answering benchmarks (SearchQA, DocVQA) fall in the p = 0.05 to 0.10 range, and on LiveMath with GPT-5.4-nano GraphSkillEvo trails SkillOpt by 0.80 percentage points. In addition, ALFWorld cells are left blank for the Codex harness because that benchmark requires persistent environment interaction not supported by the standard Codex adapter, and token consumption is not lower on every benchmark (GraphSkillEvo uses more on ALFWorld). These are open questions about scope and further validation rather than refutations of the findings.
