PersistBD lifts backdoor attack success from 20% to 74% after benign SFT and keeps it at 76% after RL
Synopsis
This work asks whether a backdoor in an attacker-supplied model survives a developer's benign supervised fine-tuning (SFT) and task-level reinforcement learning (RL) for software-engineering agents, finds that SFT substantially reduces attack success while RL often preserves or even increases the residual rate, and introduces PersistBD, which refines an already-backdoored model before release with a single LoRA adapter to raise backdoor strength and gradient compatibility, lifting attack success on Qwen2.5-Coder-7B from 20% to 74% after SFT and from 20% to 76% after SFT–RL while keeping benign task performance comparable.
Interpretation
In an SFT–RL pipeline for software-engineering agents, benign SFT substantially erodes a backdoor without eliminating it, and subsequent RL generally preserves the residual attack success and can increase it in some settings. Prior work on inherited supply-chain backdoors largely examined a single training stage; this study chains insertion, benign SFT, and task-reward GRPO into one pipeline and reports attack success at each stage. Across 3B, 7B, and 30B models, attack success is 100% after insertion and falls to 8%, 20%, and 7% after SFT; after RL it is 7%, 20%, and 8% in the random-position setting, while in the first-position setting 30B rises from 6% to 12% and 7B PersistBD rises from 34% to 42%. Each evaluation setting uses 100 triggered prefixes and 100 clean counterparts, and the false triggering rate on clean counterparts stays 0% at all evaluated RL steps.
A first-order analysis attributes backdoor survival through SFT to two factors: initial backdoor strength and the compatibility between the backdoor-loss gradient and benign training gradients. The decomposition turns the question of whether a backdoor survives into separately measurable initial strength and local gradient alignment, tested through diagnostic adapter sweeps against post-SFT attack success. Sweeping coefficients while approximately holding the other factor fixed, the weakest variants retain only 23% attack success after SFT whereas the strongest retain up to 95%; in compatibility sweeps attack success generally increases with compatibility, though the curves are not strictly monotonic. The authors note the controls isolate the factors only approximately and that the analysis is a local first-order approximation.
PersistBD refines an already-backdoored model before release with a single joint LoRA adapter that raises both backdoor strength and gradient compatibility, improving survival through the developer's SFT and RL. Unlike methods that optimize trigger tokens, PersistBD keeps the trigger fixed and instead refines the backdoored model weights, and it needs no access to the developer's SFT data, estimating the benign training direction from separate benign trajectories. Strength rises at all three scales (3B 9.5 to 11.5, 7B 10.9 to 14.3, 30B 10.0 to 16.1) and compatibility changes from negative to positive; on 7B attack success goes from 20% to 74% after SFT and from 20% to 76% after RL, on 3B from 8% to 21% and 7% to 22%, and on 30B from 7% to 37% and 8% to 36%.
At the evaluated post-training endpoints, PersistBD's benign resolved rate matches or exceeds that of the Base backdoor. The comparison separates PersistBD's additional effect from the utility cost of backdoor insertion itself: on 30B, insertion alone drops resolved tasks from 101/300 to 20/300. On 7B, PersistBD resolves 23/300 tasks after SFT versus 20/300 for Base, and both resolve 27/300 after RL; across all three models PersistBD's observed resolved rate is at least that of Base at both endpoints. Ablations show that removing the strength term drops attack success to 5%, while removing the compatibility or benign-anchor term raises the false triggering rate to 0.99 and 0.87 and drives attack success to zero.
Perspective
The results speak to developers who start from third-party weights and adapt software-engineering agents through SFT followed by task-reward RL: in this setting, pre-release refinement can markedly raise a backdoor's survival through both training stages, so defenders should detect inherited backdoors when adopting external models and throughout post-training. The method is reproducible, with code released, and the attacker needs only separate benign trajectories to estimate the benign training direction rather than the developer's SFT data.
The first-order approximation describes a single update and does not predict the full training trajectory or RL-stage behavior; the strength and compatibility sweeps isolate the factors only approximately, leaving their independent causal contributions and interactions incompletely resolved. The observed RL-stage increases in attack success are not uniform across branches, and their mechanisms and reproducibility remain unresolved. Results depend on the evaluated SFT and RL procedures and training budgets, and whether longer training, different data, or a broader range of post-training procedures can fully eliminate these backdoors remains an open question. In addition, the P-Trojan comparison and the combined-method evaluation have not been repeated on the task-instance-disjoint partitions.
