DuoOPD Rewrites Distillation Feedback from Joint Teacher-Student Outcomes, Lifting Macro Accuracy by 2.58 Points on Qwen3 and 5.98 on Llama over OPD
Synopsis
The work introduces DuoOPD, a multi-task on-policy distillation method in which the student's verified outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it—when only the teacher succeeds its verified answer becomes context for scoring the student's failed response, and when only the student succeeds a weight shared within the task reinforces the whole response, with one rule covering all four outcome combinations and no task-specific settings; across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation.
Interpretation
The paper shows that on-policy distillation (OPD) ignores correctness outcomes and therefore pushes down responses the student already answers correctly: on DuoOPD's training responses the original-context teacher-student log-ratio is negative on average even when the student is correct, while correct responses make up 64-69% of training responses. Earlier methods such as OPDVR constrain feedback direction by student correctness alone but use the teacher the same way whether or not it succeeded; DuoOPD also verifies the teacher and lets the joint outcome select the teacher's scoring context and weight allocation. The paper reports task-dependent proportions of the four teacher-student correctness combinations for Qwen3 and Llama, for example the Llama student succeeding where the teacher fails on 5.9% of biology responses but 9.8% of physics responses, together with mean log-ratio and gated-weight statistics on training responses.
DuoOPD unifies multi-task distillation under one four-outcome rule: when only the teacher succeeds its verified answer is added to the teacher's context to guide correction, when only the student succeeds all valid tokens of that response receive one positive weight shared within the task, and when both succeed or both fail teacher preferences set the strength of reinforcement or suppression. The rule conditions on each response's outcome rather than its task, so it needs no task-specific settings; weight sharing is per task rather than global because log-ratio scales differ across tasks (per-task shared weights average 0.56-0.57 on biology, chemistry, and physics, versus 0.56-0.64 and 0.57-0.64 in the two additional mixtures). Ablations show that removing teacher references lowers mean macro accuracy by 1.49 points and removing weight sharing by 0.71 points, while removing both lowers it by 2.41 points to 64.56%, close to OPDVR's 64.83%; within-task sharing outperforms across-task sharing in all three mixtures.
On biology, chemistry, and physics, DuoOPD leads all five baselines for both Qwen3 and Llama, improving mean macro accuracy over OPD by 2.58 and 5.98 percentage points and exceeding the strongest baselines, EOPD on Qwen3 and OPDVR on Llama, by 1.16 and 4.58 points. Gains appear both on questions the teacher solves and on those it fails (1.90 and 0.67 points on Qwen3, 3.86 and 2.12 on Llama), indicating the student learns more of what the teacher knows and also succeeds more often beyond it. Results average three training runs per method with reported standard deviations; evaluation samples eight responses per test question and reports avg@8, with macro being the equal-weight mean across the three tasks.
DuoOPD also leads on two more heterogeneous mixtures: 63.48% macro accuracy on scientific knowledge, understanding, and calculation, 0.86 points above the strongest baseline ExOPD, and 49.24% on physics, instruction following, and code generation, 1.41 points above OPDVR, scoring highest on every task. The paper further reports that DuoOPD attains the highest accuracy on the hardest task of every setting (physics 62.10 and 69.02, physics calculation 43.33, MBPP 32.69) without an explicit worst-task objective. Each additional mixture is trained separately with the Qwen3 pair, with tasks entering only through their verifiers; instruction following uses official IFEval prompt-level strict scoring and code uses original MBPP assertions executed in an isolated process.
Perspective
The result targets multi-task distillation settings that have a binary verifier and a teacher and student sharing a tokenizer: the paper validates on biology, chemistry, and physics knowledge, on materials knowledge, chemistry understanding, and physics calculation, and on physics, instruction following, and code generation, with teacher-student pairs Qwen3-4B to Qwen3-0.6B and Llama-3.1-8B-Instruct to Llama-3.2-3B-Instruct. The method requires caching and verifying one teacher response per training question; the paper reports this one-time cache takes 30% (Qwen) and 23% (Llama) of an OPD training loop in GPU-time, and that if the cache is shared across students, seeds, and hyperparameters the training loop plus cache overhead falls to +9.2% and +15.0%. The paper notes the rule composes with task-level methods: a verifier can validate a task-specific teacher, task weights can scale DuoOPD's per-task feedback, and joint outcomes offer a signal for task scheduling.
The paper itself lists open questions: the method relies on a binary verifier, so open-ended tasks such as writing call for feedback rules that handle graded or uncertain judgments; teacher computation could shrink further through shared caches, compressed references, and selective reference use; longer reasoning trajectories warrant study of exploration and self-correction; and each mixture is trained and evaluated on the same tasks, leaving open whether joint-outcome guidance improves transfer to held-out tasks and whether task-level outcome statistics can guide task scheduling. In addition, this reading is of the full text, but tables and figures appear here in textual form and some individual values (such as certain log-ratio and weight means) are not fully given in the text, so readers needing exact numbers should consult the original tables and appendices.
