ASPIRE uses a behavior graph and dual explore-exploit experts to find agent prompt-injection vulnerabilities, reaching 54.1% attack success on AgentDojo
Synopsis
ASPIRE is an open-ended prompt-injection red-teaming framework for LLM agents that maintains an evolving Agent Security Behavior Graph, uses complementary Explore and Exploit experts to propose consequence-centric test skeletons, has a separate realizer build concrete payloads and trigger queries, and updates the graph and cross-run strategy memory from trajectory evidence; on AgentDojo and AgentDyn it raises attack success to as high as 54.1% while covering more consequences, injection methods, environments, and behavior paths.
Figure 1 : Comparison of problem setups between ASPIRE and existing prompt-injection evaluation paradigms.
arXivInterpretation
It reframes agentic prompt-injection red-teaming from optimizing injection strings to open-ended vulnerability discovery, defining a vulnerability as an exploitable influence path in which untrusted content enters a workflow, affects intermediate decisions, and reaches an unintended action or state change. Prior benchmarks and black-box or grey-box methods typically optimize injection content after the target consequence and environment are fixed, leaving the vulnerability hypothesis predetermined; this work expands the search object to combinations of consequences, methods, environments, and behavior paths. The paper states this reframing through its problem setup and formalization, and supports it in experiments by counting each distinct consequence-mechanism-environment triple supported by at least one successful test as a unique vulnerability.
ASPIRE centers on an Agent Security Behavior Graph whose nodes capture injection environments, access tools, observations, reasoning modules, action tools, persistent states, and final outputs, with edges encoding information and control flow and each edge carrying execution evidence and a status such as untested, supported, blocked, conditional, or saturated. The graph does not reconstruct the agent's ground-truth internal computation; it represents the red-teamer's current empirical knowledge, turning what to test next into a traceable decision. Ablation shows that removing the behavior graph collapses discovery to 0.0% ASR, indicating that without a topological model of the interaction surface, LLMs cannot spontaneously find valid multi-step attack paths.
Explore and Exploit experts divide labor: the Explore expert expands the frontier toward unconfirmed consequences, under-tested environments, and uncertain paths, while the Exploit expert verifies and generalizes by completing partial trajectories, reproducing confirmed vulnerabilities, and testing transfer across paths, environments, or methods. Test skeletons are separated from concrete payload construction, with a separate realizer generating the injection payload and a benign trigger query, preventing the search from collapsing into local prompt rewriting. Explore-only discovers 81 unique skeletons at 40.0% ASR, Exploit-only reaches 45.5% ASR but only 24 unique vulnerabilities, and the dual-expert framework reaches 96 unique vulnerabilities and a 52.4% peak ASR; removing the experts drops average ASR from 52.4% to 34.7% on AgentDojo and from 37.8% to 25.2% on AgentDyn.
Cross-run strategy memory distills turning points in the search into transferable lessons, each with a category, polarity, context, claim, and action, retrieved and then retained, weakened, or suppressed according to the current graph state, failure stage, and candidate tests. Memory stores how to navigate vulnerability discovery, such as when to switch injection surfaces, repair a trigger, compare methods, or abandon a saturated route, rather than payloads to replay. Across four attacker backbones, long-term memory yields consistent positive gains, with average absolute ASR improvements of +6.9% on AgentDojo and +7.1% on AgentDyn; lightweight backbones benefit notably (Gemini-3.5-Flash-Lite at +8.1% and +8.5%).
Perspective
The work targets builder-side vulnerability discovery for LLM agents that interact with external environments, tools, and memory. Experiments run on the executable environments of AgentDojo and AgentDyn, covering banking, Slack, travel, workspace, DailyLife, GitHub, and shopping, and examine both ReAct and planner-executor architectures. Its outputs are vulnerability reports and a vulnerability map: reports reconstruct the evidence chain from injection environment to target consequence and attribute behavioral causes, while the map organizes tested combinations of consequences, methods, environments, and behavior paths, distinguishing confirmed, partial, failed, and untested hypotheses. For a reader, this offers a reusable diagnostic lens: treat injection as an influence path rather than an isolated string, and use it to assess an agent's own entry and action surfaces.
A careful reader should note that failed attempts are not treated as proof of security; their stall points only indicate where influence was interrupted and where further testing remains necessary, so blank regions of the vulnerability map are open items rather than conclusions. Coverage is uneven across benchmarks: action tool coverage is 92.6% on AgentDojo but drops to 42.9% on AgentDyn, which the text attributes to deep sequential dependencies and strict precondition constraints in dynamic environments. On diversity, weaker backbones show mode collapse in the consequence dimension (normalized entropy dropping to 0.593, Simpson index 0.475), suggesting search balance may vary with model capability. Under defenses, some domains show higher ASR for hardened agents than for the undefended baseline, and the text offers attention anchoring and signal-driven expert adaptation as possible explanations that remain to be verified. Some formulas and figures are not fully rendered in the parsed text, so exact metric definitions and algorithm details still require the original appendices.
