Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

COBRA Architecture Cuts Branch Steering Attack Success from 94.4% to 0% While Retaining 97% Benign Utility

The work systematically studies branch steering attacks on Computer Use Agents (CUAs), in which an adversary crafts untrusted data to coerce an agent down a hazardous, pre-approved branch without injecting explicit instructions, introduces STEER-Bench (101 tasks across 9 domains) showing high attack success against standard (94.4%) and vanilla Dual-LLM (89.5%) CUAs, and proposes COBRA, an architecture pairing trusted branching plans with ahead-of-time capability constraints that strictly bound the parameters and destinations each branch may execute, reducing attack success to 0% on STEER-Bench while retaining 97% benign utility.