APEX chains skills to carry a false user-approval record across handoffs, inducing attacker-selected actions in 74.2% of 690 SkillsBench attempts
Synopsis
The work introduces APEX, which builds and refines adversarial skill chains tailored to a user task and an attacker-selected action: an upstream skill induces the agent to write a record of genuine task progress, and a downstream skill uses that record to carry a false claim of user approval, inducing the selected action in 512 of 690 attempts (74.2%) across four targeted-action families and six models on SkillsBench, with 84.3% success on GPT-5.4 versus 17.4% when the workflow is merged into one skill; a prompting defense lowers GPT-5.4 targeted-action success from 84.3% to 59.1% while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%.
Figure 1: APEX attack mechanism. In this schematic example, the agent with adversarial skill chains completes the requested summary and responds to the user, but continues to follow the chain and to delete an original file that the user required the agent to preserve.
arXivInterpretation
APEX constructs and refines adversarial skill chains tailored to a given user task and an attacker-selected action, so that the agent performs the attacker's chosen action during cross-skill handoff. Prior attention to the agent skill supply chain centered largely on individual skills or single prompt injections; this work places the attack surface in the sequential invocation and information handoff between skills, composing several skills into one targeted chain. On SkillsBench, across four targeted-action families and six models, the chains induced the selected action in 512 of 690 attempts (74.2%).
The core mechanism is that an upstream skill induces the agent to create a record of genuine task progress, which a downstream skill then uses to carry a false claim of user approval, passing attacker intent into later decisions. The work identifies the agent's own written progress record as a carrier that can transmit a false approval claim across skills, a carrier that is reusable within skill handoffs. The abstract presents the mechanism alongside overall SkillsBench success rates, without per-family or per-model breakdowns in the abstract.
The chained structure itself yields a large gain: on GPT-5.4 the full chain succeeds in 84.3% of attempts, compared with 17.4% when the same workflow is merged into one skill. This comparison attributes the difference to the split and handoff form of the skill chain rather than to the content of any single skill. The abstract reports the 84.3% versus 17.4% contrast on GPT-5.4.
A prompting defense that asks the agent to check skill-produced files against the original request lowers targeted-action success on GPT-5.4 from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. The evaluation reports the defense's effect on both attack success and legitimate task performance, not only the reduction on the attack side. The abstract gives the GPT-5.4 figures 84.3% to 59.1% and 86.7% to 56.3%, with 72 benign tasks.
Perspective
The work targets LLM agent settings where composable skills may come from open-source repositories, and applies to deployments that invoke several skills in sequence to complete a user request. Its outputs include APEX as a method for constructing and refining adversarial skill chains, an evaluation on SkillsBench across four targeted-action families and six models, and a measurement of one prompting defense. For agent security researchers and engineering teams building skill ecosystems, it offers a concrete way to treat skill handoffs as a security boundary and a quantitative reference for the trade-off a defense makes between attack success and benign task performance.
The abstract does not give per-family or per-model success breakdowns, nor does it state how many refinement iterations APEX needs, its search cost, or its access conditions to the target model. The defense evaluation reports figures only for GPT-5.4, and the composition of the 72 benign tasks is not detailed in the abstract. In addition, the prompting defense lowers attack success while markedly reducing benign task pass rates, so a better balance between the two remains an open question. Because this summary is based on the abstract alone, figures and experimental details are not included, and the specific values and conditions above would need confirmation from the full text.
