Skip to main content
Back to timeline
arXivSource publication:

SynCo jointly optimizes synthesizer and reasoner in the same online batch, reaching 78.72% average across eight math reasoning benchmarks

Synopsis

SynCo formulates agentic data synthesis and task solving as a coupled multi-agent reinforcement learning problem, updating a Synthesizer and a Reasoner from the same online interaction batch: the Reasoner receives rollout-level correctness rewards while the Synthesizer receives a task-level reward gated by task quality, answer reliability, and outcome-grounded teachability, with both updates feeding the next synthesis round. Across eight mathematical reasoning benchmarks, SynCo reaches 78.72% average accuracy versus 75.85% for a frozen-Synthesizer variant and 74.96% for the strongest external baseline, ScaleQuest-Math, with gains arising mainly from problems the initial model failed to solve.

Source-provided article image: SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
Figure 1 ·

Figure 1: Overview of SynCo , a closed-loop framework for self-evolving LLMs. (1) A Synthesizer first generates capability-guided and verifiable tasks from the current learning context. (2) The resulting interaction batch is then used to jointly update the Synthesizer and the Reasoner through online multi-agent co-training. (3) The updated agents in turn reshape the future curriculum and task distribution, forming an agent–data self-evolution loop.

arXiv

Interpretation

The paper formulates agentic data synthesis and task learning as a coupled multi-agent reinforcement learning problem and proposes SynCo, in which a Synthesizer and a Reasoner have independent parameters but are updated within the same training iteration under separate group-relative policy optimization objectives. Most prior pipelines either use static corpora or generate tasks from a previous learner checkpoint and then train the learner, so the synthesis policy and task policy are not optimized together from the same online interaction; SynCo shares one online interaction batch between them. The method specifies problem formulation, capability-conditioned synthesis context, structured task cards, verified admission, agent-specific rewards, and joint updates, with reward and verification details in Appendix B.

Across eight mathematical reasoning benchmarks, SynCo reaches 78.72% average accuracy, above the frozen-Synthesizer variant at 75.85%, the strongest external baseline ScaleQuest-Math at 74.96%, and Vanilla at 72.36%, achieving best or tied-best results on six of the eight benchmarks. Relative to Frozen, gains include 2.20 points on MATH-500, 2.05 on GSM-Hard, 1.67 on SVAMP, 6.66 each on AIME24 and AIME25, and 2.57 on Minerva; the same pattern reproduces with Qwen3-4B, where the average rises from 73.72% to 75.20%. Answer accuracy is reported under a unified verification protocol, with offline synthetic-data methods, controlled GRPO variants, an iterative self-evolution baseline, and a matched frozen-Synthesizer comparison, across Qwen3-8B and Qwen3-4B.

Gains come mainly from capability acquisition rather than retention: on the 878 examples the initial Qwen3-8B failed, SynCo solves 283 (32.23%) versus 238 (27.11%) for Frozen, while on the 2,560 initially correct examples the two are 98.95% and 98.52%. 45 of the 56 additional solved examples come from the Learn subset, 80.4% of the overall advantage, indicating that jointly adapting the synthesis policy reallocates training experience toward unresolved yet learnable regions rather than merely consolidating existing competence. Benchmarks are partitioned by initial model predictions with pooled and per-benchmark decomposition, and both methods are evaluated at fixed terminal checkpoints to avoid benchmark-specific checkpoint selection.

Outcome-grounded teachability aligns strongly with tasks that produce non-degenerate group-relative feedback, while static task properties do not substitute for it: only 16.2% of the 2,940 verified tasks are gradient-bearing, yet ranking by the SynCo gated reward raises this to 62.3%, teachability alone to 64.6%, versus 12.8% for quality-only and 12.0% for novelty-only. The Synthesizer's ex-ante difficulty labels calibrate coarse difficulty (learnable group mean success rate 0.625 versus 0.484 for hard/too hard) but leave gradient-bearing rates nearly unchanged across the three groups (16.8%, 16.0%, 16.0%), and as a predictor the learnable label yields only 16.8% precision and 22.3% recall. Pooled statistics and reranking diagnostics over 2,940 verified tasks; the authors explicitly describe this as a diagnostic analysis rather than a causal training ablation, and teachability correlates with the gradient-bearing indicator at +0.934.

Perspective

The result targets mathematical reasoning training where answers can be automatically verified: the Synthesizer produces structured task cards, and a verifier checks schema completeness, problem well-posedness, answer consistency, output format, and redundancy, with an optional independent consensus check filtering tasks whose reference answers are insufficiently supported. In this setting, the paper shows that jointly optimizing the synthesis policy, at per-step cost comparable to the frozen-Synthesizer pipeline (mean update 45.1 s versus 45.8 s; median end-to-end 343.5 s versus 341.6 s), raises the productive-task rate from 13.7% to 14.5%, i.e., 4.64 versus 4.38 productive tasks per batch, and generates 3,200 task candidates over 100 steps of which 2,940 pass verification (91.9% acceptance), without human annotation or external LLM API calls. For researchers and engineering teams who want to bring data synthesis inside the training loop and have a reliable answer checker, this offers a reusable closed-loop design; for pipelines built on static corpora or offline augmentation, it highlights the time non-stationarity of task value.

Several open questions remain for a careful reader. The curriculum analysis shows that as the Reasoner improves, the gradient-bearing share falls from 31.3% in the earliest bucket to 9.5% in the final bucket while the all-correct share rises from 34.4% to 59.9%, indicating that maintaining an informative curriculum becomes increasingly difficult; the paper also notes that only 16.2% of verified tasks provide non-degenerate group-relative feedback. The signal-alignment and reranking analyses are diagnostics over an already collected task pool, which the authors explicitly say should not be read as a causal training ablation. In the efficiency comparison the two runs were executed independently, and the authors caution that wall-clock differences should be read as comparable computational profiles rather than one variant being faster. Finally, the empirical scope is limited to mathematical reasoning and two Qwen3 scales (8B and 4B); transfer to open-ended tasks, domains without automatic verification, and larger model scales remains to be tested.

Sources