CRD distills reasoning via multi-teacher cross-feedback and step-level stitching, reaching 97.3% on MATH-500 and 70.3% on AIME'25 with a 4B student trained on only 50K examples
Synopsis
The work proposes Collaborative Reasoning Distillation (CRD), in which DeepSeek-R1 and Gemini 2.5 Flash Thinking iteratively critique each other's reasoning, a dual-judge system scores each step's logical validity independently of the final answer, and a Viterbi-style coherence-aware algorithm stitches high-quality steps across teachers; students are trained with Reasoning Quality Optimization under curriculum budget constraints, and CRD-4B reaches 97.3% on MATH-500 and 70.3% on AIME'25 using only 50K training examples.
Figure 1 : Overview of the CRD Framework. Multiple teachers iteratively critique and refine each other’s reasoning through Cross-Feedback . An LLM-as-a-Judge then assigns step-wise quality scores, and the Thinking Path algorithm stitches high-quality steps across teachers while pruning flawed steps. The curated dataset is used to train students via RQO with Optimal Budget curriculum.
arXivInterpretation
Interactive multi-teacher cross-feedback has teachers review and correct each other's reasoning chains, raising candidate data quality: mathematical reasoning accuracy rises from 64.1% to 83.5% and code generation success from 52.3% to 73.9%. Prior multi-teacher distillation relied on simple aggregation or ensemble voting; this work substitutes iterative peer critique that exploits the complementary error profiles of DeepSeek-R1 and Gemini. Ablations attribute +3.4%–7.1% student gains to cross-feedback, including +7.1% on AIME'25 and +5.9% on LiveCodeBench; Appendix G shows performance converges at rounds 2–3.
Step-wise quality assessment uses two judges (JudgeLM and xVerify), two evaluations per step, and a 25% trimmed mean to assign continuous quality scores independent of the final answer, so high-quality steps inside incorrect solutions remain identifiable. Standard distillation derives signals from binary correctness or format compliance and cannot separate sound reasoning from lucky guesses; this work moves the reward signal to the individual step. The dual-judge configuration improves the student by 1.4 points on MATH-500, versus 0.6 with xVerify alone, 0.4 with JudgeLM alone, and 0.6 with GPT-4o alone; re-scoring with GPT-4o yields a mean absolute difference of 0.428 on the 0–10 scale.
Coherence-aware step stitching (Thinking Path) jointly optimizes quality and semantic coherence to build hybrid cross-teacher chains; among 2,694 tasks with incorrect pre-stitch trajectories, 528 (19.6%) become correct after stitching. Chain-level methods discard an entire solution for one flawed step, whereas this work extracts valid steps and combines them with correct steps from other teachers, producing solutions no single teacher produced alone. Ablations attribute +1.1%–3.5% to Thinking Path; the coherence-quality trade-off parameter is optimal at 0.7 (70% quality, 30% coherence), with MATH-500 dropping to 76.0% and 82.3% at the extremes.
Training students with RQO on 50K curated examples yields CRD-4B at 97.3% on MATH-500, 70.3% on AIME'25, and 57.4% on LiveCodeBench; CRD-30B-A3B surpasses its own teacher DeepSeek-R1 on math and code benchmarks. Compared with comparable methods using 220K–800K examples, this work reaches stronger performance with roughly 12 times fewer final training trajectories; the 1.5B student needs only 1,632 steps and 8.8 H100 GPU-hours. Naive SFT on the same data already reaches 90.5% on MATH-500, indicating most of the gain comes from curation, with RQO adding 1.15 points on average; under identical data, CRD leads MAGDi by 18.5% on MATH-500.
Perspective
The framework targets small-model training settings where limited compute must yield strong reasoning: students span 1.5B, 1.7B, 4B, and 30B-A3B MoE, teachers are DeepSeek-R1 and Gemini 2.5 Flash Thinking, and training uses 50K curated examples from math, code, logic, and STEM. It applies to mathematics, code, graduate-level STEM, and planning benchmarks, and is complementary to test-time scaling—the Optimal Budget curriculum (32K→10K→1K) aims to let students reason with fewer tokens, which suits latency-sensitive deployment. For teams running similar multi-teacher distillation, the reusable parts are the step-level quality scoring and coherence-aware stitching, rather than the specific teacher pair.
CRD relies on LLM teachers and judges and may inherit their biases, so plausible but incorrect steps can occasionally receive high scores; stitching selects paths that score well under the quality and coherence criteria, which does not by itself guarantee that every selected chain is logically correct. Experiments cover one teacher pair (DeepSeek-R1 and Gemini 2.5 Flash), and other teacher combinations remain to be explored; the reported training cost covers student optimization only, not teacher generation or curation. RQO is compared with naive SFT on the same data but not with DPO. The current rubric is static, limiting scalability across diverse domains, and a promising direction is to move toward learned quality models trained on preference data. In addition, several equations and the specific values in Figure 4 are absent from the parsed text, so exact hyperparameters and curve details should be checked against the original figures and tables.
