Skip to main content
Back to timeline
arXivSource publication:

Distillation attacks still steal closed-model reasoning after reinforcement learning, with summary traces matching full traces

Synopsis

The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.

AI-generated editorial illustration: Distillation Defenses Easily Break After Reinforcement Learning

Interpretation

On the same question sets, models trained with distillation followed by reinforcement learning perform best, with distillation bootstrapping subsequent RL. Prior evaluations of distillation attacks typically measured attacker performance immediately after distillation, implicitly assuming no further training; this work incorporates RL into the attack pipeline and systematically compares distillation, RL, and their combination. Comparisons across Qwen2.5 (0.5B, 1.5B, 3B, 7B, Math-7B), Llama3 (3.2-3B, 3.1-8B), and Gemma2-2B base models using the same datasets and prompts; for example, Qwen2.5-1.5B average accuracy rises from 34.5% (base) to 61.1% (distill), 64.6% (RL), and 65.1% (distill then RL).

Antidistillation sampling appears effective when evaluated after distillation, but after the student is trained with RL, the performance gap from low to mild poisoning nearly vanishes. Previous work judged defense effectiveness by post-distillation performance; this work shows that such evaluations can give a false sense of security. Using Qwen2.5-0.5B as student and Qwen2.5-1.5B-RL as teacher, low, mild, and high poisoning reduce teacher relative performance by 10%, 30%, and 86%; after RL, unpoisoned, low-poisoned, and mild-poisoned students perform similarly (about 40.6%, 40.8%, 41.3% average), while high poisoning remains lower (34.3%), which the authors attribute to an effectively much worse teacher.

Using only reasoning summaries and final answers returned by closed-source model APIs, a weak expander model can reconstruct approximate reasoning traces; distilling on these and then applying RL achieves reasoning improvements similar to using full hidden traces. Compared with prior attacks that require training an expander or extracting full hidden traces, this attack is simpler and directly targets the summary-based protection deployed by GPT, Claude, and Gemini models. In the open-source setting with Qwen2.5-14B-RL as teacher and Llama-3.2-3B-Base as student, the gap between expanded and full traces essentially disappears after RL; against Claude Sonnet 4.6, GPT-5 mini, and Gemini Flash 3.6 summaries, the first two have full-trace baselines extracted via Panfilov et al. (2026), and the simple attack recovers similar performance after RL; Gemini lacks a full-trace baseline because the extraction attack was patched.

Any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective, leading the authors to discuss limits of response-level defenses and the direction of batch-level defenses. The work extends defense evaluation from single responses to the attacker's subsequent training and to sets of related queries, noting that response-level defenses are constrained by dual-use. Based on the above experiments and multiple reported distillation attacks in 2026, the authors argue batch-level defenses could cover a broader class of vulnerabilities, while real-time batch-level detection remains a nontrivial open problem.

Perspective

The work targets researchers studying distillation attacks and defenses, security teams at model providers, and institutions shaping AI safety policy. Its conclusions apply to settings where attackers can obtain reasoning summaries and final answers through APIs and can perform subsequent RL training; the authors argue realistic attackers likely use distillation followed by RL, so defenses should be evaluated under this setting. The authors also propose batch-level defenses as a more promising direction and note that domain-specific response-level safeguards (such as restricting cybersecurity queries) are sensible within a given domain.

All experiments use models with at most 3B parameters, whereas realistic distillation attacks may use models with orders of magnitude more parameters; experiments focus on math reasoning, leaving other verifiable tasks such as long-horizon agentic coding to be tested; the proposed attack is unoptimized and the authors do not consider it optimal; Gemini lacks a full-trace baseline, so conclusions about it are inferred; extraction of full traces from closed-source models relied on endpoints that were unpatched at the time, and those conditions may change over time.

Sources