Skip to main content
Back to timeline
arXivSource publication:

The Same Zero: A VAL-Tier Framework Shows Identical ASR Hides Different Guarantees, as a VAL-Guided Stack Holds 0.000 Attack Success at 1.000 Benign Success

Related research and updates

Synopsis

The authors apply Verification Autonomy Levels (VAL, L0–L5) to 22 LLM-agent security defenses and run an equal-budget comparison of a VAL-guided stack (confirmation gate + schema sandbox) against a mainstream intuition stack (prompt hardening + keyword filter) across 50 scenarios, 12 attack variants, and adaptive/white-box/PAIR escalation, finding the VAL stack holds 0.000 attack success at 1.000 benign success while the intuition stack reaches 0.000 ASR but kills all benign actions, and that across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0→1.9%→6.2%) while the VAL stack holds within its ODD (0→0→0), its only breach a disclosed out-of-ODD password gap (0.5%).

AI-generated editorial illustration: The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security

Interpretation

The paper introduces and applies Verification Autonomy Levels (VAL), classifying defenses by where their guarantee comes from: L0 (LLM self-declaration), L1 (deterministic rules), L2 (objective ground truth), L3/L4 (decidable completeness), L5 (impossible), mapped onto 22 agent-security defenses. The defense landscape was already dense (prompt hardening, content filters, permission gates, sandboxes), but no framework told a deployer what a defense actually guarantees or where that guarantee comes from; VAL supplies that attribution layer. The abstract reports the taxonomy is falsifiable, with 10/10 prediction hits on frozen cards (flagged).

The paper reports the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) reaches 0.000 attack success at 1.000 benign success. The comparison is against a mainstream intuition stack (prompt hardening + keyword filter), which also reaches 0.000 ASR but kills all benign actions, i.e., security by model-behavior luck rather than structure. 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation, ~7,000 testbed calls plus ~10,000 harness calls on AgentDojo/JADE; on AgentDojo banking the VAL stack shows 0.5% ASR at 79.7% utility versus 4.3% undefended.

Across testbeds of rising attack-surface hardness, the intuition stack's zero drifts while the VAL stack holds within its operational design domain (ODD). This directly supports the paper's core claim that zero is an outcome, not a guarantee, and that identical ASR can imply different guarantees. The intuition stack drifts 0→1.9%→6.2% (n=16 on JADE), while the VAL stack is 0→0→0, its only breach a disclosed out-of-ODD password gap (0.5%).

Perspective

The work targets engineers and security evaluators making deployment decisions for LLM-agent defenses, within the operational design domain defined by the tested scenarios, attack variants, and testbeds (including AgentDojo and JADE); VAL tiers serve as an attribution tool for inventorying where a defense's guarantee comes from, and the VAL-guided stack's conclusions are bounded by its ODD, with the paper itself disclosing an out-of-ODD password gap.

Readers should still watch: how VAL tiering covers defenses beyond the 22 examined; how the intuition-stack drift and VAL-stack stability behave on more testbeds and longer attack escalation; how the ODD boundary is defined and extended given the disclosed out-of-ODD password gap (0.5%); and what utility-versus-ASR trade-off thresholds are acceptable in different business settings. The abstract does not give full per-scenario detail, so the exact distribution of numbers still requires the original text.

Sources