Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

The Same Zero: A VAL-Tier Framework Shows Identical ASR Hides Different Guarantees, as a VAL-Guided Stack Holds 0.000 Attack Success at 1.000 Benign Success

The authors apply Verification Autonomy Levels (VAL, L0–L5) to 22 LLM-agent security defenses and run an equal-budget comparison of a VAL-guided stack (confirmation gate + schema sandbox) against a mainstream intuition stack (prompt hardening + keyword filter) across 50 scenarios, 12 attack variants, and adaptive/white-box/PAIR escalation, finding the VAL stack holds 0.000 attack success at 1.000 benign success while the intuition stack reaches 0.000 ASR but kills all benign actions, and that across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0→1.9%→6.2%) while the VAL stack holds within its ODD (0→0→0), its only breach a disclosed out-of-ODD password gap (0.5%).