Skip to main content
Back to timeline
arXivSource publication:

SkillDRE evolves malicious agent skills through a dual-stage loop of pre-execution scanning and runtime feedback, reaching a 45.28% average attack success rate across four victim models on SkillsBench

Synopsis

SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.

AI-generated editorial illustration: SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Interpretation

The paper proposes treating feedback from both the pre-execution scanning stage and the runtime defense stage as learning signals, evolving complete malicious skill packages through a dual-stage closed loop rather than optimizing against a single defense stage. Prior work on adversarial skill evolution typically considers detection or execution outcomes at one stage; SkillDRE returns every runtime-guided revision to the pre-execution stage for rescanning and further optimization before re-execution, chaining the two stages into a cross-stage loop. The abstract describes the framework as fully automated and reports evaluation on SkillsBench across four victim models, with an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%.

Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then keeps both fixed during evolution. Automating the construction and validation of the malicious objective and judge rule gives skill-implementation evolution a stable evaluation anchor, while requiring preservation of legitimate task capability. The abstract explicitly describes this automated construction and validation process, the fixing of objective and judge rule, and the preservation of benign task capability; the concrete construction details are not expanded in the abstract.

The final submitted skills achieve attack effectiveness while receiving no SkillScan findings and largely preserving benign-task performance. This indicates attack capability and detectability under pre-execution scanning can be optimized together rather than traded against benign functionality. The abstract reports that final submitted skills receive no SkillScan findings and largely preserve benign-task performance; the degree of preservation and the evaluation protocol are not given in the abstract.

The paper argues that evaluating either defense stage in isolation can miss the resulting attack capability. This observation makes the evaluation protocol itself part of the conclusion, suggesting that looking only at pre-execution scanning or only at runtime defense is insufficient to characterize the actual risk after skill evolution. The abstract states this as the interpretation of the results, supported by the 45.28% average attack success rate and the 40.3% improvement over the strongest baseline.

Perspective

The work targets the agent-skill-package form and applies to skill ecosystems that deploy both pre-execution scanning and runtime defense; its conclusions rest on evaluation with SkillsBench and four victim models, with the main evidence being a 45.28% average attack success rate, a 40.3% improvement over the strongest baseline, zero SkillScan findings for final skills, and largely preserved benign-task performance. For defenders, this means the feedback from both stages should be examined within a single evaluation loop rather than scored separately; for red-team research, it offers a reusable automated pipeline for probing the actual attack capability of evolved skills while preserving benign functionality.

The abstract does not give the per-victim-model result distribution, the task and sample scale of SkillsBench, the composition of the baseline set, or how the degree of benign-task performance preservation is measured; it also does not state which scanner version and configuration the zero SkillScan findings correspond to. In addition, under which task conditions the automated construction of the malicious objective and judge rule remains reliable, and how stable the loop is when defenders keep updating, are questions left open at the abstract level.

Sources