HASTE evolves agent harnesses automatically from sparse threat evidence, lowering attack success rates while preserving benign-task utility across backbone models and attack types
Related research and updatesSynopsis
HASTE is a multi-agent framework that evolves agent harnesses from sparse threat evidence, such as brief descriptions or a few attack examples in threat reports and preprints, through an adversarial interplay between safety-specification generation and attack-case generation, feeding evaluation outcomes back into both processes so harnesses evolve against emerging attacks beyond the initially observed evidence; experiments across multiple backbone models, attack types, and evidence forms show it consistently reduces attack success rates while preserving benign-task utility.
Figure 1: Illustration of harness evolution from a full benchmark versus sparse threat evidence . (a) With a full benchmark, diagnosis and proposal directly revise the harness based on judge feedback. (b) With sparse threat evidence, attack cases and a safety specification are generated to guide harness adaptation and iteratively optimized using judge feedback.
arXivInterpretation
It introduces HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence. Harness adaptation has relied on manual work, while rapidly emerging attacks outpace manual adaptation; HASTE automates this process and targets the sparse-signal setting where threat reports and preprints offer only brief descriptions or a few attack examples. The abstract describes the framework as an adversarial interplay between safety-specification generation and attack-case generation and provides a code link; implementation details, evidence scale, and experimental setup are not expanded in the abstract.
Safety-specification generation guides harness updates toward identified safety vulnerabilities, while attack-case generation probes for remaining safety vulnerabilities after each update. The two processes are not a one-way pipeline but an adversarial loop: specifications indicate what to fix, and attack cases test what remains after each fix, enabling harness evolution against emerging attacks beyond the initially observed evidence. The abstract presents this adversarial interplay as a mechanism description and states that evaluation outcomes are fed back into both processes; no quantitative ablation or component-contribution results are given.
Experiments across multiple backbone models, attack types, and evidence forms show HASTE consistently reduces attack success rates while preserving benign-task utility. The results span multiple backbone models, attack types, and evidence forms, indicating the approach is not tied to a single model or attack shape, and that it lowers attack success rates while maintaining normal task performance. The abstract reports an overall consistency conclusion across conditions without specific attack success rate values, utility metrics, sample sizes, or statistical tests.
Perspective
The work targets developers and security researchers who need to continuously harden safety constraints in agent harnesses, in settings where only sparse threat evidence is available, such as brief descriptions or a few attack examples in threat reports or preprints. Its goal is to keep evolving harnesses beyond the initially observed evidence while balancing reduced attack success rates with preserved benign-task utility across multiple backbone models, attack types, and evidence forms. The abstract also provides a code link, supporting reproduction and extension under the same setting.
The abstract does not disclose specific attack success rate values, how benign-task utility is measured, the scale of evidence, the list of backbone models, or the list of attack types, nor whether evaluation includes statistical tests or ablations; these details affect judgments about effect magnitude and component contributions. In addition, this reading scope is the abstract only, without the body, figures, or appendix, so full validation conditions for the described mechanism and conclusions still require consulting the original text.
