Skip to main content
Back to timeline
arXivSource publication:

SEAD recasts tool-agent attack and defense as partially observed state control: DART lifts semantic attack success by 18.8–35.9 points, SAGE cuts executable attack success from 48.0% to 4.0%

Synopsis

The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.

AI-generated editorial illustration: SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents

Interpretation

The paper formalizes agentic safety as partially observed state control in a shared execution loop: the attacker supplies instructions, the target chooses concrete tool actions, and the defender decides before an action takes effect, with harm defined by the state execution reaches rather than by a single response. Prior attack work largely constructs attacks at the dialogue level and prior defenses start from role-specific objectives; here the design requirements for both roles are derived jointly from persistent effects, role-specific partial observations, and execution feedback. The paper gives formal definitions of state transitions, observation projections, a harmful-state predicate, and the harm-enabling transition, and argues that when visible history is compatible with two states requiring opposite decisions, history alone cannot determine the boundary crossing.

DART decomposes harmful goals into locally plausible steps and runs tree search over executed trajectory prefixes, using real tool-execution feedback plus an executable verifier and two LLM-based heuristic scores to value nodes. Compared with direct requests, fixed multi-turn decompositions (MTA), synthetic tool-use context (STAC), and conversational intent hijacking, DART lets realized execution rather than intended steps drive the next instruction and branch selection. On 187 tasks across four target models, DART has the highest attack success under both criteria, exceeding each target's best baseline by 18.8–35.9 points semantically and 8.1–17.9 points on executable checks; a case shows it expanding an alternative initial prefix after repeated target refusals to reach confirmed success.

SAGE can actively acquire environment-state evidence through read-only, defender-private queries before allowing or blocking each pending action, and re-gates every action proposed after a block. Relative to decision-only binary defenders and released checkpoints such as StepGuard, TS-Guard, and Safiron, SAGE treats investigation itself as part of the safety mechanism rather than relying on visible history alone. On recorded trajectories from a 75-task subset, SAGE reaches 95.79% benign all-pass, 34.55% exact interception, and 92.73% interception, with the highest harmonic scores among the listed defenders; versus the fine-tuned binary control it gives up 4.21 points of benign utility while gaining 65.45 points of harmful interception.

Online attack–defense evaluation shows that blocking changes the subsequent trajectory, so protection must be measured on the continuation after intervention; SAGE lowers success under four attack methods and transfers to PostgreSQL and Web, which were excluded from defender fine-tuning. The paper separates fixed-trajectory evaluation, which isolates a single decision, from online evaluation, which lets the attacker revise plans in response to refusal feedback and tests whether defense prevents rather than merely delays harm. Online, SAGE reduces DART's semantic success from 65.3% to 10.7% and executable success from 48.0% to 4.0%; on 30 PostgreSQL/Web tasks it reduces semantic successes from 20 to 3 and executable successes from 9 to 1; a case shows an earlier refusal does not guarantee final protection.

Perspective

The framework targets tool-mediated execution in which consequential actions pass through a pre-execution gate and state investigation uses read-only, defender-private queries; experiments span four tool domains (Filesystem, Terminal, PostgreSQL, Web), attack effectiveness is evaluated on four target models, defense evaluation uses GPT-5.6 Luna as the target, and PostgreSQL and Web are held out of defender fine-tuning. For readers building or evaluating pre-execution gating, it offers a reusable data construction (controlled initial states, replayable tool environments, task-specific executable checks) and two complementary measurement perspectives: benign preservation and interception timing on fixed trajectories, and online success rates on the continuation after intervention.

Concurrent processes, compound actions, and disclosure through assistant output are listed as directions requiring extensions to the intervention interface; online defense evaluation centers on a single target model, and broader target coverage plus controlled studies of individual search and investigation components are named as future work. In addition, the loaded text is the full paper with appendices but presents figures as tables, so readers wanting to check details such as the distribution of block timing against the original figures would still need the source.

Sources