Public articles linked to the same research event.
arXiv HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.
HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.
HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.
HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.