SAKI routes teacher supervision through maximal-coupling accept/correct events, lifting Mean@8 and Pass@8 for both 1.7B and 0.6B students across seven math reasoning benchmarks
Synopsis
SAKI realizes a KL-constrained teacher-guided rollout through maximal coupling and reuses the realized accept/correct events as a token-level supervision router: accepted positions keep sampled-token reverse-KL, correction positions switch to direct supervision on the teacher's highest-probability token, and the correction probability is exactly TV(p_t,q_t) so the same trust-region radius upper-bounds intervention frequency; an engine-resident speculative verifier preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x, and across seven mathematical reasoning benchmarks SAKI improves Mean@8 and Pass@8 over the matched teacher-guided baseline for both 1.7B and 0.6B students.
Interpretation
SAKI turns the accept/correct events produced by maximal coupling into the supervision router itself: accepted positions retain sampled-token reverse-KL, while correction positions switch to negative log-likelihood on the teacher's Top-1 token. Earlier teacher-guided rollouts such as TRB change where supervision is queried but keep the same per-prefix reverse-KL objective; SAKI exploits the coupling events that arise naturally when realizing that guided distribution, decoupling where to supervise from what to supervise without an extra token-selection heuristic or routing threshold. The paper proves Proposition 1 and Proposition 2: the correction probability is exactly TV(p_t,q_t) and is the minimum intervention probability among couplings with these marginals; Corollary 1 shows the trust-region radius both constrains rollout deviation and upper-bounds intervention and specialized-supervision frequency. Appendix J.2 reports observed correction probability staying below the Pinsker upper bound and vanishing as the radius reaches zero.
Across seven mathematical reasoning benchmarks, SAKI achieves the best macro Mean@8 and Pass@8 at both the 1.7B and 0.6B student scales. Relative to TRB, which shares the same exact-q rollout and annealing schedule, the 1.7B student improves from 27.9/44.6 to 29.0/47.5 and the 0.6B student from 17.2/33.6 to 18.4/35.6; gains over SKD are larger. Mean@8 improves over TRB in 13 of 14 student-benchmark pairs. Students are Qwen3-0.6B-Base and Qwen3-1.7B-Base distilled from the same Qwen3-4B-Base-GRPO teacher on DAPO-Math-17K with a 200-step budget and a rollout batch of 64 prompts with eight responses each; evaluation covers MATH-500, HMMT-Feb26, AIME 2026, AIME 2025, AMC 2023, Minerva Math, and OlympiadBench with eight samples at temperature 1.0 and maximum response length 16,384.
The placement of correction-triggered supervision carries information: under an equal teacher-mode supervision budget it outperforms both random and TV-weighted placement. Random-TM reaches 28.30 Mean@8 / 45.50 Pass@8 and TV-Weighted-TM reaches 28.24 / 46.77, while SAKI reaches 29.00 / 47.50, indicating that scalar conflict measures help but realized coupling corrections provide a stronger routing signal. Both controls match the number of teacher-mode updates per exact-q trajectory; for TV-Weighted-TM the correction mask determines only the per-trajectory supervision budget, not the selected locations. A fixed-prefix probe of 2,048 positions from 213 prompts shows the advantage growing from the lowest to the highest conflict quartile, with a Q4-Q1 difference of 0.87 to 1.29 (95% CI), carried mainly by teacher-supported non-argmax tokens.
Engine-resident speculative block verification preserves the exact-q trajectory distribution and coupling semantics while raising matched-workload rollout throughput from 776 to 3,276 tokens/s. That is a 4.22x speedup over the external-loop implementation and over 42% of the 7,760 tokens/s theoretical ceiling set by unguided student-only generation; Appendix K's implementation history shows a token-wise prototype at 5 tokens/s and an early external block path at 165-299 tokens/s. Proposition 3 proves that if proposal prefixes match sequential maximal coupling, only the prefix up to the first rejection is committed, the correction is sampled from the exact residual distribution, and later speculative states are discarded, the generated trajectory follows the same autoregressive distribution as sequential sampling; Appendix K.5 validates bridge construction, coupling, training state, and routing.
Perspective
The result targets on-policy distillation validated on mathematical reasoning: students are Qwen3-0.6B-Base and Qwen3-1.7B-Base, the teacher is the same Qwen3-4B-Base-GRPO, training uses DAPO-Math-17K with a 200-step budget, and evaluation covers MATH-500, HMMT-Feb26, AIME 2026, AIME 2025, AMC 2023, Minerva Math, and OlympiadBench. The method itself does not depend on mathematical content, so it can be applied to other distillation pipelines that need online teacher queries, and the engine-resident speculative verifier offers a reusable execution structure for exact-q rollouts. In the main training schedule correction supervision is transient: once the trust-region radius is annealed to zero the method returns to student-rollout reverse-KL training, so the mechanism is designed as an early-training guide.
Correction supervision stops once the radius is annealed to zero; the fixed-prefix probe shows alignment gains still visible after 149 reverse-KL-only steps, but how long that persistence holds under longer training and larger models remains open. Placement controls and fixed-prefix analysis are run mainly on the 1.7B student, so whether the 0.6B scale shows the same conflict-adaptive pattern is not yet confirmed. Throughput numbers depend on response length, correction rate, and committed tokens per verification wave, so the speedup may shift under different workloads. In addition, this evidence bundle is a full-text parse in which some specific percentage-point values appear as placeholders, so individual magnitudes can only be described from the prose rather than checked digit by digit.
