Skip to main content
Back to timeline
arXivSource publication:

SAIL reaches top comparison scores on SciCode and ArxivDIGESTables with 35B total and 3B active parameters

Synopsis

The SAIL team introduces SAIL, an open scientific agent model with 35B total and 3B active parameters, developed through a science-aware improvement loop in which frontier-model agents diagnose its task failures and build training tasks from papers and code repositories, then trained via supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning; it improves over its base model on all 13 evaluations and achieves the highest SciCode and ArxivDIGESTables scores in the comparison.

Source-provided article image: SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
Figure 1 ·

Figure 1 : Performance across scientific benchmarks.

arXiv

Interpretation

SAIL is an open model for literature research, scientific coding, and multi-step research workflows, with 35B total and 3B active parameters, and the model plus most of its training data are released. Compared with open-weight models of similar scale, it covers literature retrieval, code execution, and end-to-end research rather than a single benchmark. It improves over the base model Qwen3.6-35B-A3B across all 13 evaluation settings, with the largest gains on LitQA-search (+37.33 points), E2E-Bench Basic (+30.71), CORE-Hard (+24.37), and E2E-Bench Hard (+21.21); against the compared 35B models it leads on 12 settings, with LitQA2-FullText the exception.

A science-aware improvement loop has frontier-model agents analyze SAIL's failed trajectories, identify the missing scientific judgment, and construct problems, demonstrations, and executable tasks with environments from paper collections and scientific GitHub repositories. Task topics and forms are chosen from the model's diagnosed capability gaps, so the same paper or code resources can serve different training objectives as weaknesses change. Diagnosis covers the exploration-selection tradeoff in literature search, scientific assumptions and reasoning in coding, and planning and revision in longer investigations; generated material is validated against its sources and task-specific checks, including execution where applicable.

A training recipe and infrastructure consolidate these tasks into a single model: SFT initialization, specialist training, multi-teacher on-policy distillation (MOPD), and agentic reinforcement learning, sharing one on-policy trajectory collection and execution system. MOPD applies token-level specialist supervision on the student's own trajectories, including tool observations and compacted context, while agentic RL then trains multi-step execution from task feedback, with both sharing the rollout and collection pipeline. The framework is built on verl, separating the training engine from the rollout engine, and uses a token-in/token-out interface recording input and output token IDs, rollout log probabilities, and training masks; long trajectories are split into segments at context compaction while retaining the same episode identity, with rewards and advantages aggregated over the original trajectory.

Evaluation shows SAIL achieves the highest SciCode and ArxivDIGESTables scores in the comparison and approaches or exceeds substantially larger open-weight models on several tasks. With 35B total parameters it exceeds the 744B GLM-5.2 on SciCode and ranks second on the unweighted mean aggregate of twelve task scores at 59.76. SciCode is 50.35 versus GLM-5.2's 47.57; ArxivDIGESTables is 35.24 versus DeepSeek-V4-Flash-0731's 35.13; ScholarQA-CS2 is 86.51 versus GLM-5.2's 87.87; the aggregate is 59.76 versus 59.15 for DeepSeek-V4-Flash-0731 and 63.30 for GLM-5.2.

Perspective

The work targets researchers and engineering teams working on literature retrieval and synthesis, scientific coding, and multi-step research workflows involving tool use, especially those deploying open scientific agents under a limited parameter budget. The improvement loop presumes access to failure trajectories, usable paper collections and scientific code repositories, and the tools and environments required for executable tasks; evaluation is conducted in the same benchmark environments, with each model using its recommended inference settings.

Next development cycles depend on the coverage of the resource collections and the quality of diagnosis and feedback; the paper notes that scientific assumptions and experimental interpretations remain difficult to check through execution alone. Results are presented in text and tables, while the contents of Figures 2 through 6 are not expanded in the text, so details of the failure-diagnosis process and system architecture still require the original figures. Full-text question answering has a weaker relative ranking (91.67, a 3.57-point gap to the leader), the strongest larger models retain a lead on repository execution and long research workflows, and DiscoveryBench shows a smaller improvement over the base model.

Sources