8B multimodal verifier SciGen-Verifier tops Qwen3.8-Max with 86.28% judgement accuracy on a scientific image verification benchmark and serves as an online critic for iterative correction
Synopsis
The work builds SciGen-Verify, a benchmark for explainable verification of scientific image generation (1,350 samples: 406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge, with a three-tier cascading protocol over binary judgement, explanation, and editing instruction), and develops SciGen-Verifier on Qwen3-VL-8B-Instruct trained by cold-start supervised fine-tuning (48,311 instances) followed by a curriculum-based two-stage reinforcement learning pipeline (rubric-guided process reward, then outcome reward), reaching 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy on the benchmark while also serving as an online critic for iterative image rectification.
Interpretation
The paper introduces SciGen-Verify, described by the authors as the first benchmark for explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge under a three-tier cascading protocol over binary judgement, explanation, and editing instruction. Prior verification benchmarks and reward models target natural images and focus on generic image-text alignment and human aesthetic preference; existing scientific image generation benchmarks introduce fine-grained rubrics but serve merely as static, offline testbeds, lacking a dedicated verifier that actively produces explainable, actionable correction feedback. The benchmark contains 1,350 samples (406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge), built through AI-assisted annotation plus independent review by 10 domain experts (two graduate students each for mathematics, physics, chemistry, biology, and geography), with samples retained only when both experts approve; negatives come from controlled visual and textual perturbation strategies and pass a filtering stage that discards severe blurriness, illegible text, or clear generation failures.
The paper develops SciGen-Verifier-8B, initialized from Qwen3-VL-8B-Instruct and trained by cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline: stage one uses a rubric-guided process reward to strengthen scientific reasoning exploration, and stage two uses an outcome reward to align judgement, explanation, and editing instruction. The authors observe that SFT failures predominantly stem from incomplete evidence exploration rather than incoherent reasoning, so process-level rewards first broaden reasoning coverage before outcome-level rewards are introduced, avoiding unstable training from applying outcome rewards too early. The SFT corpus has 48,311 instances (19,678 instruction following, 20,367 multidisciplinary reasoning, 8,266 world knowledge), kept via rejection sampling where the binary judgement matches ground truth and the explanation is validated; RL uses GRPO, with stage one retaining 15K partially successful queries and stage two using 685 fully annotated examples with answer, explanation, and editing-instruction supervision; training and benchmark are split at the source-sample level so a source problem and all its perturbed variants belong to exactly one set.
On SciGen-Verify, SciGen-Verifier-8B attains 86.28% overall judgement accuracy, 81.79% explanation accuracy, and 57.42% editing-instruction accuracy, surpassing the strongest baseline Qwen3.8-Max and larger proprietary models such as Gemini-3.5-Flash and GPT-6-Astra. The authors report that this 8B backbone generalizes robustly across all three domains, with world knowledge reaching 94.21% judgement, 91.53% explanation, and 65.48% editing-instruction accuracy, eclipsing Qwen3.8-Max by 6.45, 4.41, and 5.67 points respectively; OmniVerifier-7B, trained for generic image verification, performs poorly on the benchmark, indicating that scientific verification extends beyond generic text-image alignment. Comparison covers 15 leading multimodal models, all evaluated under the same input-output protocol and the same judge models; explanations and editing instructions are judged by GLM-4.7 and Doubao-Seed-2.0-Lite, whose agreement with human experts on 300 labeled instances is 95.3% and 94.7% against a human-human ceiling of 98.3%, and re-evaluation with GPT-5-mini preserves the model ranking.
Ablations show that explicit chain-of-thought and the curriculum-based two-stage RL are jointly indispensable, and SciGen-Verifier can act as an online critic driving iterative test-time scaling to progressively correct scientific image generation errors. Direct outcome prediction without CoT reaches a competitive 81.0% judgement accuracy but collapses on faithful feedback (explanation 65.9%, editing instruction 36.4%), and applying RL to that direct-prediction paradigm brings negligible gains; skipping process-level grounding caps reasoning accuracy at 84.9% and editing-instruction accuracy at 51.2%. In the test-time scaling application, an initial image from Wan2.6-t2i is judged by the verifier and triggers regeneration or local editing when needed, for at most 10 steps, lifting the relaxed score from 58.7 to 63.5; for reference, Gemini-3.5-Flash and Qwen3.8-Max reach 64.9 and 64.4 after ten steps, only 1.4 and 0.9 points higher, while Seed-2.0-Pro is more volatile and ends at 62.6.
Perspective
The result targets verification of scientific image generation and editing, suited to educational and generative settings that must decide whether an image satisfies multi-step scientific constraints and then supply actionable correction instructions; the benchmark covers instruction following, multidisciplinary reasoning, and world knowledge, and the model is built on Qwen3-VL-8B-Instruct trained on a 48,311-instance corpus split from the benchmark at the source-sample level, so its scope is a dedicated verifier and online critic for this task rather than a general image-quality evaluator.
Editing-instruction accuracy is markedly lower than judgement and explanation accuracy across all models, and the authors identify synthesizing actionable corrective edits as a major bottleneck, so further progress on this step is worth watching; the benchmark and training corpus share upstream sources, and although they are split at the source-sample level, cross-source generalization remains an open question; the relaxed score used in test-time scaling is not directly comparable to the benchmark's three hierarchical metrics, and its iterative trajectory does not rise monotonically at every step.
