Skip to main content
Back to timeline
arXivSource publication:

OmniVBench and Omni-R2V: A Shared Foundation for Evaluating and Training Omni Reference-to-Video Generation

Synopsis

This work introduces the OmniVBench benchmark and the Omni-R2V dataset, evaluating omni reference-to-video generation with 7 task families, 18 fine-grained tasks, 813 evaluation cases and 12,172 factor-grounded checklist items, and releasing 340K processed training samples; evaluation of several open- and closed-source models reveals clear performance gaps across task families and evaluation dimensions.

AI-generated editorial illustration: OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Interpretation

It introduces OmniVBench, expanding reference-to-video evaluation from a content-centric focus to five single-reference dimensions (content, motion, style, structure, narrative) plus multi-content and cross-aspect multi-reference settings, totaling 7 task families, 18 fine-grained tasks and 813 evaluation cases. Whereas existing benchmarks such as OpenS2V-Eval, VACE-Bench, UniVBench, IntelligentVBench and FashionVideoBench mainly cover content references, this benchmark marks content, motion, style, structure, narrative and multi-reference categories together in its comparison table. The task taxonomy and case counts are listed item by item in the main text and Appendix Table A.1 (e.g., content 112, motion 95, style 43, structure 108, narrative 120, multi-reference 335), and all samples were manually verified by three annotators.

It proposes a factor-grounded evaluation protocol that decomposes each case into case-specific checklists explicitly assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction, totaling 12,172 checklist items. Existing evaluation largely relies on fixed similarity metrics or holistic VLM judgments of reference consistency, whereas this protocol separates Reference Fidelity, Instruction Realization and Video Quality, and subdivides fidelity into content, structure, motion, style and narrative sub-dimensions. Human evaluation on 100 cases covering all 7 task families and 18 sub-tasks, yielding 965 model outputs, shows automatic evaluation agreeing with human judgments at Pearson 0.81/0.78/0.86 and Spearman 0.77/0.74/0.82 for RF, IR and VQ; the factor-grounded checklist reaches Spearman 0.77 and 0.74 on RF and IR, above holistic evaluation at 0.67 and 0.69.

It constructs and releases the Omni-R2V dataset with 339,570 (about 340K) processed training samples covering 7 task families and both image and video reference modalities, together with task-specific reference-target pairing and instruction generation pipelines. Existing datasets such as OpenS2V-5M, Phantom-Data and MuSS center on image references and content tasks, whereas this dataset covers content, motion, style, structure, narrative, multi-content and cross-aspect tasks in its comparison table and directly provides processed reference-video pairs. The data draws primarily on an in-house professional video corpus supplemented with public data, passing through source-video processing, task-specific pair construction, candidate filtering and instruction generation; clip durations extend to 20 seconds and resolutions reach 1080p and 2160p+; human spot checks of roughly 100 samples per task family report average pass rates of 95.4%, 97.0% and 94.2% on three criteria.

Evaluation of 11 advanced open- and closed-source models finds that no model performs consistently well across all task families, with content reference scores generally higher while larger gaps emerge on motion, style, structure, narrative and multi-reference settings. The evaluation explicitly separates Reference Fidelity, Instruction Realization and Video Quality, and further splits out reference-factor disentanglement and routing versus target compliance, exposing differences that holistic evaluation tends to blur. Results are reported per task family and overall in tables, e.g., among open-source models MiniMax H3 at 72.41 overall and Bernini at 56.95, and among closed-source models Seedance 2.5 at 72.68 and Gemini Omni at 70.76; several models score relatively high on target compliance while showing substantially lower disentanglement and routing scores, especially on multi-content and cross-aspect tasks.

Perspective

The work targets research and engineering settings in reference-to-video generation: on the evaluation side it suits model comparisons that need to diagnose reference-factor preservation, disentanglement and routing, and multi-reference composition; on the data side it suits training conditioned on image or video references in single- or multi-reference compositions. The benchmark and dataset are non-overlapping, so they can serve evaluation and training separately, and the task-specific pairing pipelines and checklist protocol can transfer to other reference-conditioned generation tasks.

Checklist evaluation relies on a VLM as judge, whose stability across task families warrants continued observation; the video quality dimension is computed by DOVER++, Aesthetic Predictor V2.5 and UnifiedReward 2.0 separately and is independent of reference fidelity and instruction realization, leaving room to discuss how the three jointly explain model capability; in addition, some references in data construction are synthesized by generative models, and how the difference between such references and real captured references affects training is a question worth tracking.

Sources