Skip to main content
Back to timeline
arXivSource publication:

Meituan LongCat team lifts Qwen3 math reasoning by 1.67–2.75 points using teacher parameter-neighborhood perturbations

Synopsis

The work introduces Neighborhood OPSD (N-OPSD): in on-policy self-distillation, local parameter perturbations of the reference-conditioned teacher yield a set of frozen experts; offline greedy selection builds a complementary pool by rewarding filtered reference-token gains beyond the pool's current best, and online routing uses MaxPeak to pick an anchor token and a quantile rule to pick the level of support among anchor-consistent experts, improving the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, with only the distilled student needed at inference.

AI-generated editorial illustration: Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

Interpretation

The authors find that a privileged teacher's reference-aligned supervision is not confined to its unperturbed parameter setting: experts obtained by local parameter perturbations supply complementary corrections at different reference positions, and their pool covers more reference positions than the unperturbed M0. Prior OPSD used one fixed teacher parameter setting at every student-visited state, and multi-teacher distillation typically assigned teachers by domain, task, or task accuracy; here diversity comes from nearby parameter settings interpreting the same privileged reference context. On 500 held-out questions at σ = .002, the pool covers 26.53% of reference-side SCGate-retained positions versus 15.69% for M0, and pool coverage is higher at every tested radius; coverage requires positive filtered PTA from at least one expert.

Offline greedy selection adds experts by rewarding filtered reference-token gains beyond the pool's current best, outperforming random selection, ranking by answer accuracy, greedy sample coverage, and independent PTA Top-K. Expert choice shifts from 'who is more accurate' to 'who adds reference-aligned corrections the current pool does not already supply', explicitly accounting for overlap in corrections within the pool. On Qwen3-8B with routing and gating held fixed, the full method reaches 66.57 Average@12 versus 65.65 for independent PTA Top-K, 65.37 for greedy sample coverage, 65.00 for Quality Top-K, and 64.29 for random Top-K; the authors report gains on all three benchmarks.

Online routing separates the anchor direction from its level of support: MaxPeak takes the token with the highest probability assigned by any expert as the anchor, then a quantile rule selects among experts whose top token matches it, and the student learns the chosen expert's full next-token distribution. Unlike averaging expert distributions, majority-vote consensus, or always taking the highest-peak expert, this rule lets a single expert propose the direction while avoiding clipping of the anchor's forward-KL term when the teacher-student probability gap is large. With the same pool, quantile routing at q = .75 beats uniform averaging by 1.94 points, token consensus by 1.57, and random routing by 1.66; accuracy rises from q = 0 to q = .75 and then drops 1.29 points at q = 1, indicating the highest-peak expert need not be the best training target.

The benefit of the selected supervision source persists when the divergence or target distribution changes, and inference needs only the single distilled student. It separates 'expanding the supervision source' from 'adapting the divergence or loss weights', showing routed targets can be combined with JSD, RKL, or compressed targets. With pool and router fixed, N-OPSD outperforms OPSD+SCGate under both JSD and RKL; even with a binary Top-1-tail soft target it still exceeds OPSD+SCGate by 2.12 points, and the full-vocabulary target scores 1.30 and 1.20 points above the two compressed variants.

Perspective

The result targets on-policy self-distillation for mathematical reasoning with a privileged teacher that sees a reference solution, and applies to the Qwen3-1.7B, 4B, and 8B family evaluated on AIME 2024, AIME 2025, and HMMT February 2025. The authors' defaults are K = 25 experts, routing quantile q = .75, online SCGate threshold τ = .99, selection clip κsel = κ = .06, with calibrated radii of .0006, .0012, and .002 for 1.7B, 4B, and 8B. The expert pool stays frozen during training, only student parameters receive gradients, and inference uses only the distilled student, so downstream users need not keep any perturbation experts in serving. The authors also note that offline selection uses filtered PTA on reference prefixes while online routing happens at student-visited states, so the pipeline is best read as a two-stage 'build the team offline, pick the member per state online' scheme that can be extended to other model families and to code generation, multi-hop QA, and tool use.

The expert pool stays frozen throughout training while the student keeps changing, and the authors themselves raise the possibility that the experts most useful to the student may differ from those selected offline, suggesting periodic PTA recomputation and reselection as future work while weighing the cost of repeatedly evaluating all candidates. On radius, reference-side coverage and anchor accuracy keep rising at larger radii, yet student-prefix continuation accuracy and downstream average fall after σ = .002, so reference-side metrics are not a direct proxy for student supervision quality and the authors calibrate the radius on student prefixes instead. Pool size and the best tested quantile are coupled: the best tested quantile is .75 at K = 25 and .50 at K = 50, and the best K = 50 result still falls below K = 25. The evidence bundle is the full paper, but some analysis details (the complete reference-side audit and the per-item continuation probe) appear in appendix tables, so judgments about those intermediate quantities still require checking the appendices.

Sources