SelfEvoSkill revises agent skills via paired execution audits without external outcome supervision, reaching 73.9% mean reward across 89 tasks
Synopsis
The work introduces SelfEvoSkill, which pairs executions with and without a skill while holding agent, task, and workspace fixed, links the resulting behavioral differences to relevant skill passages, and iteratively revises and reevaluates those passages, with checks derived once from the task's explicit requirements preventing constraint violations; across 89 containerized tasks in 8 professional domains under the same best-of-two evaluation budget, it attains 73.9% mean task reward, exceeding the no-skill condition by +33.0 percentage points and the benchmark's curated skill by +17.2 percentage points, while skill evolution receives neither task rewards nor benchmark verifier outputs.
Figure 1: Three paradigms for skill revision. Outcome-Gated (left) and Deployment-Feedback (center) rely on external outcome signals. SelfEvoSkill (right) uses behavioral differences between matched executions with and without the skill to locate relevant skill passages and iteratively revise them, without external outcome supervision.
arXivInterpretation
It proposes paired execution audits as a substitute for external outcome supervision in skill revision: holding agent, task, and workspace fixed, it compares behavior between a with-skill and a without-skill run, links the differences to specific skill passages, and iteratively revises and reevaluates those passages in subsequent executions. Existing skill-revision methods typically rely on external outcome supervision such as task rewards, performance on a separate validation set, or deployment feedback to decide which updates to retain; this work instead uses execution behavior itself as the signal, targeting settings where such supervision is costly or unavailable when adapting a skill to a new task. The abstract states that skill evolution receives neither task rewards nor benchmark verifier outputs, so its signal source is explicitly distinguished from external outcome supervision by design.
Revisions are constrained by checks derived once from the task's explicit requirements, preventing revised skills from violating task constraints. Task constraints are fixed up front as checks rather than re-derived from outcome feedback each round, giving the revision process a constraint basis even without outcome supervision. The abstract describes this as 'Checks derived once from the task's explicit requirements,' a method-design statement.
Across 89 containerized tasks in 8 professional domains under the same best-of-two evaluation budget, mean task reward is 73.9%, which is +33.0 percentage points over the no-skill condition and +17.2 percentage points over the benchmark's curated skill. The comparison provides both a no-skill baseline and the benchmark's curated skill as references, and emphasizes that skill evolution used neither task rewards nor verifier outputs, indicating the gain comes from behavior-difference-driven revision rather than outcome supervision. The abstract reports task count, domain count, evaluation budget, and two percentage-point deltas against controls, making this a quantitative result with comparisons.
Ablations show separate contributions from paired comparison and repeated revision: using two with-skill runs instead of a matched with/without-skill pair reduces mean reward by 4.6 percentage points, and limiting revision to one round reduces it by 9.0 points relative to repeated revision. These two ablations isolate the role of the paired signal construction and of the iterative revision mechanism, rather than reporting only the overall gain. The abstract reports both deltas as controlled ablations with explicit numerical values and directions.
Perspective
The result is meant for settings where an agent skill must be adapted without external outcome supervision, especially where task requirements can be stated explicitly and executions can be paired reproducibly in a containerized environment; for such readers it offers an actionable path to locate skill passages from behavioral differences and revise them iteratively, with quantitative references against no-skill and the benchmark's curated skill under the same best-of-two budget.
The abstract does not state how the 89 tasks are distributed across the 8 domains, how reward is defined, or the statistical uncertainty, nor does it give per-domain or per-task results; the concrete implementation of holding agent, task, and workspace fixed during paired execution, the process for deriving checks from task requirements, and the trade-off between revision rounds and cost all need confirmation in the main text. The abstract also does not report a direct comparison against methods that rely on external outcome supervision under the same budget, so the boundary between behavioral-difference signals and outcome supervision remains an open question.
