Skip to main content
Back to timeline
arXivSource publication:

SRFT lets LLM agents learn from their own failures, cutting AgentDojo attack success from 7.59% to 1.26%

Synopsis

The work proposes Self-Reflection Fine-Tuning (SRFT): it injects attack instructions into clean expert trajectories, samples the agent's own potentially failing actions under the attacked context, and has an expert model generate structured self-reflection reasoning that contrasts unsafe and optimal actions, which then supervises fine-tuning; instantiated as SR-Agent on Llama-3.1-8B-Instruct and Qwen3-8B, it reduces AgentDojo attack success rate from 7.59% to 1.26% and from 16.97% to 1.05% respectively, and keeps a low attack success rate under the RL-Hammer adaptive attack.

Source-provided article image: Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience
Figure 1 ·

Figure 1: Left: Supervised fine-tuning learns from fixed defensive patterns, which may become ineffective when encountering unseen attacks. Right: Self-reflection fine-tuning learns from failure experiences, enabling it to handle a broader range of attacks and identify previously unseen ones.

arXiv

Interpretation

SRFT turns prompt-injection defense from imitating expert actions into learning from the agent's own failure experience: it injects attack instructions into clean expert trajectories, then rolls out the target agent under the attacked context to expose failure cases where it follows the injection. Compared with prior alignment methods that rely on static expert trajectories or preference optimization (StruQ, SecAlign, Meta-SecAlign), the framework exposes the agent to its own real failure behavior under adversarial conditions rather than only imitating predefined defensive patterns. Training data comes from 5 TOUCAN platforms: 402 user tasks, 38 malicious injection tasks, 3698 trajectories and 22339 assistant steps, with 38% of tool observations carrying an injection; 3 responses are sampled per assistant step.

A three-part self-reflection generated by an expert model (summary of context and original goal, identification of the injection and its malicious intent, safety-aware action analysis with counterfactual consequence prediction) serves as supervision, teaching the agent to explain why the optimal action beats hijacked ones. The reflection contrasts the agent's own sampled actions, providing positive justification for the optimal action and negative evidence against unsafe candidates, unlike defenses that only suppress patterns or rank preferences. Reflections are generated by Claude Sonnet 4.6 conditioned on injection ground truth and the reference agent's failures, averaging 396 tokens (median 401); the training objective factorizes into generating the reflection and then the action.

On AgentDojo, SRFT lowers attack success while raising task utility: for Llama-3.1-8B-Instruct, ASR drops from 7.59% to 1.26%, Benign Utility rises from 27.84% to 37.11%, and Utility under Attack rises from 23.08% to 29.08%. Against Meta-SecAlign-8B on the same backbone, SR-Agent achieves a better utility-security trade-off without sacrificing general capability (comparable to base on MMLU, MMLU-Pro, IFEval and BBH). AgentDojo spans four environments (Banking, Slack, Travel, Workspace) with 949 trajectories; the Qwen3 family reproduces the ASR reduction at both 4B and 8B scales.

Under the RL-Hammer adaptive attack, SR-Agent keeps final ASR below 20% while Meta-SecAlign exceeds 60%; ablations show removing failure-experience analysis raises final ASR from 17.0% to 65.0%, and disabling thinking at inference raises AgentDojo ASR from 1.05% to 14.33%. The results indicate robustness comes not from memorizing safe responses but from explicit reasoning and action analysis grounded in the agent's own failures, offering an experience-driven path against unseen attacks. RL-Hammer trains a separate attacker per target model; the InjecAgent 510 direct-harm cases are split into 310 training, 100 validation and 100 test. On Llama, final ASR is 98.5% for the base, 69.5% for Meta-SecAlign and 17.5% for SR-Agent.

Perspective

The result targets defense of tool-augmented multi-step agents against indirect prompt injection, in settings where external observations such as webpages, emails, documents, retrieved content or API responses are corrupted; training data covers five platforms (hotel booking, email sender, Windows command line, Markdown downloader, Minecraft Wiki), and the attacker is assumed to have access only to realistic attack surfaces. For researchers and engineering teams seeking to improve agent security without sacrificing general capability, SRFT offers a reusable training recipe and open-source code; at inference its reflective supervision relies only on the dialogue itself, needing neither injection ground truth nor the reference agent's failure samples.

Worth watching: the reflective supervision depends on expert-generated injection ground truth and the reference agent's failure samples, neither of which is visible at inference, so the model must reproduce the three-part analysis from the dialogue alone; the training set contains only 38 malicious injection tasks with triggers drawn from a pool of 750 templates, so the coverage of attack forms remains an open question. In addition, AgentDojo uses the default important_instructions attack and the RL-Hammer evaluation performs no checkpoint selection, so behavior under other attack configurations remains to be characterized; extending the paradigm to computer-use or web agents is left as future work.

Sources