APEX moves indirect prompt-injection defense to the execution boundary, cutting attack success to 0% on five of six benchmarks and 0.56% on the sixth
Related research and updatesSynopsis
The work presents APEX, which shifts indirect prompt-injection defense from recognizing attack patterns to checking each proposed effect at the execution boundary: before untrusted execution it compiles the user's authorization into a contract, WRAP admits an effect only when the contract and runtime evidence reconstruct its authorization, and PLANT places probes on the contract's endorsed dependencies so unendorsed use of runtime information reveals itself before commitment; across six benchmarks and 13 baselines APEX reaches 0% attack success on five benchmarks and 0.56% on the sixth, and holds 0% under three adaptive attacks spanning all three capability-unit types.
Figure 1: Left: injections arrive through many carriers and compositions, but become consequential only at the execution boundary, where they appear as unauthorized actions or unendorsed uses of information. APEX actively checks each proposed effect there. Right: security ( 1 − ASR 1-\mathrm{ASR} ) against attack utility on one benchmark per capability unit; APEX (star) sits at the top right of each panel, and attains the lowest ASR with high utility on all six benchmarks, 0 % 0\% on five.
arXivInterpretation
The paper decomposes boundary risk by task endorsement into two classes: an effect that is unauthorized, and an authorized effect that draws on runtime information the task never endorsed; both checks are brought to a single execution boundary. Prior defenses each hinge on one signal, such as content separability, alignment of behavior with the task, rules written in advance, provenance recovered from the run, or decoys placed beforehand; APEX instead makes task authorization the sole criterion, so it does not grow with the attack carrier. The decomposition is stated in the threat model and formalized in Appendix B as an authorization-soundness lemma and a task-level theorem, and is supported experimentally by six benchmarks against 13 baselines.
WRAP traces the contract backward from a proposed effect and resolves it forward from runtime evidence, admitting an action and each argument only when the contract and receipts reconstruct them, without propagating security labels. Unlike policy-based defenses, the contract states not only which effects are permitted but along which dependencies their arguments may be determined; unlike taint-style defenses, it does not require labels to survive the model's own rewriting. Ablation shows that removing WRAP raises ASR on ASB-OPI from 0% to 57.60%, close to the undefended 59.80%, and raises it to 4.15% on MCPTox and 2.22% on SkillInject.
PLANT places observation, dependency, and substrate probes carrying episode-unique tokens on the contract's endorsed dependencies; a token appearing in a proposed argument or released output constitutes evidence of unendorsed use. Probe placement is derived from the contract rather than decided heuristically, and a probe replaces the identity-bearing value itself, so carrying it forward carries the token while dropping it leaves the injection without a handle. At least one probe is deployed in 80.1% of the 5,486 attack cases, and 779 of 10,193 deployed probes (7.6%) produced a gating witness; PLANT produced no false positive across all cases.
Safe continuation rebuilds state from proved bindings after a rejection, first attempting deterministic repair, then handing a sanitized state to a fresh agent for replanning, and otherwise aborting the branch, without ever modifying the contract. Recovery is performed by a controller that is deliberately not an agent, replaying proved values and invalidating evidence, so it changes the execution path without expanding task-authorized authority. Removing continuation lowers attack utility from 85.05 to 62.20 on ASB-OPI and from 75.52 to 48.15 on MCPTox while providing little security benefit; on ASB-OPI one replan returned the agent to the user task in all but 11 of 2,040 cases.
Perspective
The result applies to externally consequential operations mediated through registered interfaces, covering Tools, MCP servers, and Skills and their compositions; bypassing effects or state changes lie outside the guarantee. A defender must first compile an authorization contract faithful to the user's request, and each capability unit must be registered once with its interface, I/O schemas, external effects, and probeable surfaces. For deployers who want uniform mediation of capability invocations without capability-specific security logic, this boundary check is directly reusable; tasks that are long-horizon, or that intentionally delegate decisions to external content, first require explicit limits on who may supply instructions and which actions they may authorize.
Contract fidelity is the premise of protection strength: compilation runs before any untrusted content is read, so a misstated contract blocks legitimate work rather than admitting harm, but ambiguous or underspecified requests may yield an over-restrictive contract. PLANT's coverage rests on deployment and probe preservation, and the share of cases admitting a placement is markedly lower on MSB (41.2%) and SCR-Bench (37.5%) than on the other benchmarks, with WRAP alone constraining carriers where no behavior-preserving placement exists. The residual risk is harm whose effect and information dependencies are both task-endorsed, as in the single SkillInject case behind 0.56%: the attack stays inside an authorized effect and carries no identity that can travel, calling for a finer specification of intent or a task-specific safety policy. Long-horizon tasks extend in theory, but accurate compilation of complex objectives and conditional branches still needs dedicated long-horizon safety benchmarks.
