OmniEdu fine-tunes 4B/9B/27B models on 69,999 capability-oriented examples, beating each base model across curriculum grounding, K–12 problem solving, and pedagogical tutoring
Synopsis
Researchers from Peking University, the University of the Chinese Academy of Sciences, and Zhongguancun Academy built OmniEdu, an open family of K–12 learning-and-teaching foundation models, trained with a capability-oriented instruction corpus organized around subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding (69,999 examples and 15.96M supervised response tokens, of which 60,951 examples and about 12.0M tokens are education-specific), fine-tuning 4B, 9B, and 27B backbones with full-parameter supervised fine-tuning and reporting consistent gains at every scale on curriculum-grounding, K–12 problem-solving, and pedagogical-tutoring benchmarks, with OmniEdu-27B reaching 63.12% EM / 76.69% F1 on K12-Bench, 85.89% on MathFish, 86.
Interpretation
Introduces OmniEdu, an open model family at 4B, 9B, and 27B scales whose educational supervision is organized around four capabilities: subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding. Prior educational models often specialize in either subject problem solving or tutoring, and their training mixtures are commonly organized by source or task without explicitly balancing these capabilities; this work makes the capability taxonomy the shared language of both data construction and evaluation. The paper defines the capabilities, maps sources to categories, and trains three model scales, comparing each against its corresponding base model under each benchmark's official protocol and metric.
Builds a reproducible capability-oriented data pipeline — deterministic cleaning and evaluation decontamination, LLM-assisted semantic auditing and rewriting, task-specific quality scoring, token-budgeted diversity selection, and pedagogical instruction assignment — reducing an education-specific pool of roughly 1.34M examples to 60,951 high-quality examples (about 12.0M supervised response tokens), plus 9,048 general instruction examples, for 69,999 examples and 15.96M tokens in total. Rather than mixing data by source or subject alone, the pipeline assigns each example a single primary capability and one of 20 task-specific pedagogical system instructions, separating an example's content from the response behavior it is meant to teach. The paper reports stage-by-stage counts (870,711 → 440,100 → 121,318 → 60,951) and names the tools used: Qwen3.5-122B-A10B-FP8 for semantic auditing, GPT-5.6-Terra for fine-grained scoring, and k-center greedy over BGE-M3 embeddings for diversity selection.
Cross-scale evaluation shows education-oriented tuning improves curriculum grounding, K–12 problem solving, and pedagogical tutoring over the corresponding base model at every scale. The gains are not limited to answer accuracy: K12-Bench EM rises from 52.11% to 63.12% at 27B, EDUMATH MaC improves by 20.80, 15.40, and 16.35 points at 4B, 9B, and 27B, MathTutorBench Scaffold win rate improves by 55.37 and 61.26 points at 4B and 9B, and LongTutor Evidence average rises from 36.80% to 78.20% at 27B. Results are presented as paired base-versus-tuned comparisons across K12-Bench, MathFish, EDUMATH, GAOKAO-Bench, EXAMS-V, MDK12-Bench, MathTutorBench, TutorBench, and LongTutor, with full breakdown tables in the appendix.
Uses IFEval, GPQA, and MMMU-Pro as auxiliary diagnostics and reports that educational specialization does not come with a broad loss of general capability. Treats the cost of specialization as an auditable comparison rather than reporting educational gains alone. MMMU-Pro overall accuracy rises from 50.46% to 52.60% at 4B, 58.38% to 60.75% at 9B, and 64.97% to 67.98% at 27B; the paper notes that changes vary by MMMU-Pro discipline.
Perspective
The work targets K–12 learning and teaching, with model scales of 4B, 9B, and 27B trained by full-parameter supervised fine-tuning (LLaMA-Factory, maximum sequence length 32,768, 3 epochs). It is meant for builders of educational assistants that need curriculum localization, error diagnosis, and scaffolded tutoring, and for researchers who want to reuse its data pipeline; the project website, dataset, and three model checkpoints are released. Evaluation is limited to the curriculum-grounding, K–12 problem-solving, and pedagogical-tutoring benchmarks listed in the paper, plus the three general diagnostics IFEval, GPQA, and MMMU-Pro.
The paper notes that knowledge-state diagnosis remains comparatively difficult, with OmniEdu-27B and OmniEdu-9B reaching diagnosis accuracies of 54.04% and 53.55%, and states that these "relatively low accuracies indicate that knowledge-state diagnosis remains comparatively challenging." MMMU-Pro changes vary by discipline, and the paper reports only the overall improvement. In addition, several open-weight educational baselines do not support image inputs and receive only the textual component for image-dependent examples, a setting worth keeping in mind when comparing across models. The semantic auditing and fine-grained scoring stages depend on specific large models (Qwen3.5-122B-A10B-FP8 and GPT-5.6-Terra), so reproducibility depends on their availability. The main text defers complete breakdowns to the appendix, so subset-level details require checking the appendix rather than the summary tables.
