Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

HASTE evolves agent harnesses automatically from sparse threat evidence, lowering attack success rates while preserving benign-task utility across backbone models and attack types

HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.