A safety-aligned but backdoored teacher passes its backdoor to a clean student through on-policy distillation, reaching up to about 70% trigger attack success at a 3% poisoning rate
Synopsis
This work studies whether on-policy distillation (OPD) for safety can propagate backdoors: the authors build safety-aligned but backdoored teachers with bi-GRPO and distill initially clean students through OPD, finding that a 3% poisoning rate yields up to about 70% trigger attack success, that more training epochs (67% after 16 epochs with only 10 poisoned samples) and top-K KL both accelerate backdoor transfer, and that a KL-reward-clipping method, Lazy Defense, delays this transfer.
Figure 1: On-policy distillation with poisoned data and backdoored teacher.
arXivInterpretation
On-policy distillation can transfer a teacher's hidden backdoor to an initially clean student even at low poisoning rates. Prior OPD and OPD-for-safety work typically assumes the teacher and training data are trustworthy; this work systematically examines the safety risk when the teacher carries a backdoor. Teacher-student pairs are built on Qwen2.5-1.5B/3B/7B in same-model and cross-size settings, with 1,000 clean training samples and poisoning rates of 1%-10%, evaluated on DAN, DNA, Addition, StrongREJECT, and AdvBench using Llama Guard 3; in the same-model setting, 3% poisoning gives Trigger ASR of about 67% (3B) and 50% (7B), with Clean ASR below 5%.
Increasing training epochs substantially raises backdoor transfer with very few poisoned samples, while safety on trigger-free inputs stays largely unchanged. Under a fixed training budget, Trigger ASR was low at low poisoning rates; this work shows that this may reflect insufficient training rather than an inability to learn the backdoor. In the same-model setting with only 10 poisoned samples (1% poisoning), Trigger ASR rises from near zero at epoch 8 to about 67% (3B) and 27% (7B) at epoch 16; in cross-size OPD, 3% poisoning rises from near zero at epoch 8 to about 81% at epoch 16, and 5% poisoning from near zero at epoch 4 to about 87% at epoch 16, with Clean ASR remaining low throughout.
Top-K KL accelerates backdoor transfer compared with sampled-token KL. Prior work found that varying K had little effect on overall training performance; this work shows that the choice of KL estimator significantly changes backdoor learning dynamics. With Qwen2.5-3B distilling 1.5B at 5% poisoning, Trigger ASR under top-K KL rises several epochs before a similar increase appears with sampled-token KL; the authors explain the delay with a token-level gradient decomposition (direct and softmax-normalization indirect components) and with Jaccard similarity of student-teacher token overlap rising over training.
Lazy Defense, which clips KL rewards to make student updates less aggressive, delays backdoor transfer in low poisoning rate settings while largely preserving trigger-free safety. The method requires no knowledge of the trigger or which samples are poisoned and applies the same reward clipping to all samples, making it a mitigation that does not rely on poison identification. Compared with top-K KL and sampled-token KL across four Qwen2.5 teacher-student pairs at poisoning rates of 1%-10%; for the 3B student at 3% and 5% poisoning, and for another pair at 3% poisoning, Lazy Defense keeps Trigger ASR low for more epochs.
Perspective
The results apply to OPD safety-alignment pipelines that use external teachers and potentially poisoned training data, particularly same-model and cross-size Qwen2.5 teacher-student pairs at 1%-10% poisoning rates under fixed or extended (up to 16 epochs) training. For practitioners, this suggests that when adopting OPD for safety alignment, the trustworthiness of teacher and data sources, the number of training epochs, and the choice of KL estimator should all be part of the risk assessment, and that reward-clipping approaches such as Lazy Defense can be tried without knowledge of the trigger. For researchers, it provides an analytical frame linking backdoor transfer to token-level gradients and student-teacher token overlap, which can inform more reliable defenses.
Readers should still watch: the findings are based on bi-GRPO-constructed teachers and the Qwen2.5 family, so transfer magnitude under other backdoor injection methods or model families remains unclear; Lazy Defense delays rather than prevents backdoor transfer, and the authors note that reliably preventing backdoor inheritance while preserving safety benefits remains an open challenge; moreover, the threat model assumes users do not know the trigger or which samples are poisoned, and how to define the trust boundary for teachers and data in real deployment needs further study.
