Skip to main content
Back to timeline
arXivSource publication:

Two-stage gated auditing cuts LLM audits from 8.15 to 2.33 per run at 72.8% detection on OpenAgentSafety

Synopsis

The work presents a two-stage LLM agent monitoring framework in which single-event and trace-sequence gate rules select pending actions for inspection and an LLM audit agent judges each selected action in the context of the inspected agent's preceding trace before execution, with an optimizer LLM jointly refining gate rules and audit prompts. On the OpenAgentSafety test split, the tuned gated audit reduces mean audits from 8.15 to 2.33 per run and tokens from 47.8k to 14.6k, with 72.8% detection versus 81.5% for auditing every action and a lower false-positive rate (5.9% versus 11.8%); in two simulated multi-agent scenarios it flags all attack runs while cutting audit token usage by more than 80%.

Source-provided article image: Efficient Auditing of Adversarial AI Agent Behavior from Agent Traces
Fig. 1 ·

Fig. 1: Motivating example. (a) The attack succeeds without monitoring. (b) The attack is detected and blocked before completion with monitoring.

arXiv

Interpretation

A two-stage monitoring framework in which a deterministic gate first selects pending actions and an LLM audit agent then judges each selected action before execution using the agent's preceding trace, returning benign, suspicious, or malicious verdicts; flagged agents have their pending action blocked, the agent terminated, and an incident report retained. Relative to predefined-rule guardrails and to auditing every action with an LLM, the design confines semantic inspection to gate-selected actions, preserving semantic judgment while compressing audit volume. On the OpenAgentSafety test split (149 runs across 72 unseen tasks), the tuned gated configuration uses 2.33 audits and 14.6k tokens per run versus 8.15 and 47.8k for full auditing, with 72.8% versus 81.5% detection and 5.9% versus 11.8% false-positive rate.

Gate rules and audit prompts are jointly refined by an optimizer LLM from observed development failures rather than hand-written from a fixed security policy; the objective balances true-positive rate, false-positive rate, and token cost relative to the seed configuration, and the optimizer is instructed to describe only general behaviors and structures without copying file names or specific values from failure cases. Adapting rules and prompts through data-driven iteration lets the framework cover cross-event patterns that a single record cannot reveal, such as sensitive data entering context followed by data egress. The tuning trajectory raises true-positive rate from 47.5% to 95.0% while reducing mean audits by 62% and tokens by 54%, with false-positive rate rising only from 5.0% to 7.5%; all five revisions were accepted for improving the objective.

In two directly run simulated multi-agent scenarios (an external attacker inducing vault-key disclosure and a compromised internal agent inducing a destructive datastore command), gated auditing and full auditing both detect all attacks at the execution boundary of the harmful action with no false positives on benign runs. This complements the public-benchmark evaluation, which is limited to replaying released traces, by providing end-to-end evidence from multi-agent systems the authors run themselves. Ten attack runs and ten benign runs per scenario; gated auditing reduces token usage by 82% in Scenario 1 and 86% in Scenario 2, with the gate triggering once in Scenario 1 and not at all in Scenario 2.

The work reports provider safety-guardrail refusals and their effect on metrics: 15 of 347 requests (4.3%) were rejected for the tuned gated audit and 16 of 1,215 (1.3%) for full auditing, with refusals attributed to harmful content in the trace; refusals are treated as detections, and ignoring them would lower the tuned gated audit's true-positive rate to at least 55.6%. Refusals are reported as a distinct phenomenon with a recommended fail-closed deployment policy rather than bypassing provider safeguards. Refusal counts and their distribution across runs (14 unsafe and one safe run for the gated audit) come from the experimental record; the post-ignoring figure is labeled a lower bound because auditing stops after a refusal, leaving subsequent verdicts unobserved.

Perspective

The framework targets agent deployments whose traces can be logged across three surfaces (operational, cognitive, contextual), with logs stored append-only in an isolated environment agents cannot access, and it assumes attackers cannot tamper with or falsify trace provenance; trace tampering is explicitly outside scope. It covers both a white-box setting (internal attacker whose prompts and reasoning can be logged) and a black-box setting (external attacker whose interactions can be logged but whose internal prompts and reasoning are unavailable). For operators seeking to reduce per-action LLM auditing cost, the result offers a reproducible path: select with a deterministic gate, semantically audit the selected actions, and on a flag block the pending action, terminate the agent, and retain an incident report. The authors note that the reductions in false alarms and token consumption are especially valuable in deployments dominated by benign traces, provided the drop from 81.5% to 72.8% detection is acceptable.

Tuning raises the validation-split false-positive rate from 12.1% to 24.2% while the test split shows 5.9%; the authors attribute part of this to 32 of the validation set's 51 tasks also appearing in the training split and to benign runs that repeatedly elicit suspicious verdicts, so how the flagging rule behaves across task distributions remains to be seen. Ignoring provider refusals would lower the tuned gated audit's true-positive rate to at least 55.6%, and that value is a lower bound because auditing stops after a refusal, leaving subsequent verdicts unobserved. Removing the suspicion-accumulation criterion would reduce true-positive rates to 61.7% for full auditing and 63.0% for the tuned gated audit, showing that the flagging rule definition directly sets the detection-versus-false-alarm tradeoff. The authors list future directions including benign runs with legitimate sensitive operations, more and more diverse scenarios, controlled trace-generation models for more direct comparisons, and stronger monitor-aware attackers that avoid explicit cues. In addition, the public-benchmark evaluation replays released trajectories rather than re-running agents.

Sources