AgentTracer traces indirect prompt injection as intent drift, reaching 94.17% injection-point accuracy
Synopsis
The work presents AgentTracer, an intent-aware tracing framework that treats indirect prompt injection as task intent drift: it recovers implicit decision dependencies among tool calls to build an Intent-Driven Execution Graph, combines the user request with an operation knowledge base to construct a user intent authorization space, and traces backward from an anomalous tool call to reconstruct the attack chain. On AgentDyn and InjecAgent logs combined with a historical-log pool of 1,800 normal requests and 9,000 background tool calls without IPI, it achieves 94.17% injection-point accuracy and 93.56% path precision, improving injection-point accuracy over existing methods by 18–54%.
Figure 1 . An Example of an Indirect Prompt Injection (IPI) Attack. The malicious instruction induces an attack chain interleaved with user-authorized tool calls. An injected release-page instruction causes SSH-key registration, email retrieval, account verification, and memory clearing. Existing tracing strategies recover incomplete or noisy explanations.
arXivInterpretation
Formulates IPI tracing as fine-grained alignment between task intent and tool calls, and introduces the Intent-Driven Execution Graph (IntentExecGraph) that connects dispersed tool calls by task intent using decision dependencies. Existing tracing methods rely mainly on explicit control-flow and data-flow dependencies, whereas attacker-intent-driven tool calls may exchange no data or share no parameters; this work instead uses contextual decision relations extracted from reasoning records to establish connections. In the ablation study, replacing the IntentExecGraph with a temporal graph that connects adjacent calls by execution order reduces injection-point accuracy by 17.00% and path coverage by 31.78%, indicating a substantive contribution of this intermediate representation.
Proposes a knowledge-enhanced user intent alignment method that combines the user request with an operation knowledge base to construct a structured user intent authorization space (UA-Scope), evaluating tool calls along action, target, and constraint dimensions to identify intent drift. User requests typically state only high-level goals and omit execution details and authorization boundaries; the operation knowledge base supplies operation boundaries, prerequisites, follow-ups, targets, and checkpoints, accommodating legitimate supporting operations without extending the authorized scope. In the ablation study, replacing the UA-Scope and alignment predicates with direct LLM judgments drops injection-point accuracy from 91.00% to 64.00%; on 394 evaluable execution edges from 100 injection-free logs, the edge-level false-positive rate is 3.52%.
Designs event-guided pruning and backward tracing that narrows the tracing space at the analysis, graph, and path levels, tracing Misaligned edges backward from the anomalous event to reconstruct the attack chain and locate the injection source and injection point. Historical logs contain many unrelated requests and calls, and building a global graph would introduce noise and cost; this work anchors on the anomalous event and retains only the relevant log segment and intent-drift path. Removing event-guided pruning in favor of full-history LLM analysis yields tracing effectiveness close to the full method but increases average analysis time by 30.98% and token consumption by 51.24%, showing that pruning reduces cost while preserving effectiveness.
Constructs and manually annotates datasets of agent execution logs covering diverse scenarios and multi-step IPI attacks, and evaluates end-to-end and comparatively on a historical-log pool with normal-task noise. The evaluation setup includes both interleaving of legitimate and malicious calls within a single execution and long-accumulated noise from normal requests and background calls, with a background-to-attack-chain tool-call ratio of approximately 3,000:1. In the end-to-end experiment, 300 logs are randomly selected from each dataset, yielding an average IPI detection rate of 97.34%, injection-point accuracy of 94.17%, path coverage of 90.25%, and path precision of 93.56%; in the comparison experiment, injection-point accuracy exceeds all baselines by 18.00–54.00%.
Perspective
The method targets agents whose execution logs contain user requests, reasoning records, and tool records (tool calls and tool results), and is best suited to agents following ReAct or a similar perceive–plan–act–observe loop, because such agents record reasoning before tool calls, providing evidence for extracting contextual decision relations. It applies to settings where attackers can embed natural-language instructions in external resources processed by the agent but cannot modify the original user request or the execution logs, and the analysis focuses on IPI attacks in which malicious instructions affect reasoning and lead to observable tool calls. Attacks that do not result in tool execution, and code-level injection that does not surface as separate agent tool calls, are not directly covered by the current design; the paper notes that system-level tracing techniques based on system calls, file accesses, and network activity could be integrated to link execution records to the agent tool call that launched the code, extending the attack chain to code-level behavior. The design can also be extended to persistent cross-session IPI: when the root node of an intent-drift path is a persistent entity, the tool call that created or modified it can be found in historical logs, and the earlier-session graph can be connected to the current graph through the shared entity.
Tracing errors arise mainly from two factors: IntentExecGraph construction depends on extracting contextual decision relations from natural-language reasoning records, and the LLM may extract incorrect source entities, target entities, or intended operations, while different aliases for an entity may cause a correctly extracted relation to be grounded in the wrong entity node. In addition, an agent may deviate from the user request even without IPI; such genuine execution deviations are correctly labeled Misaligned but may form additional intent-drift branches unrelated to the attack chain. The exploratory reproduction of cross-session persistent IPI (45 cases, injection-point accuracy above 85%) is excluded from the formal evaluation because the sample size and covered scenarios are insufficient to support a general conclusion, leaving broader evaluation of persistent entities and cross-session tracing to future work. Furthermore, some comparison methods have not released complete implementations, and the paper reproduces them following published algorithms, prompts, and rules, using the same dataset, anomalous events, annotations, and LLM backbone to reduce differences introduced by the experimental setup.
