Skip to main content
Back to timeline
arXivSource publication:

Ovis-Embedding puts text, image, video and audio into one retrieval space on a shared Qwen-Omni backbone, scoring 58.46 overall on MMEB-v3 at 3B scale

Synopsis

The work reports the Ovis-Embedding family of universal omni-modal embedders: it uses a pretrained Qwen-Omni model as a shared backbone, removes the speech-generation pathway and takes the final-layer hidden state at the last non-padding token as the embedding, trains on an omni-modal corpus of roughly 50M query-target pairs with homogeneous-source sampling, focal loss and embedding distillation, and adds low-rank feature decomposition for elastic embedding dimensions, reporting leading results on MMEB-v3, MMEB-v2, MVEB, MAEB and RTEB.

AI-generated editorial illustration: Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Interpretation

Ovis-Embedding-Omni-3B reaches an overall score of 58.46 on MMEB-v3, 5.19 points above the strongest baseline Tianmu-Emb-Uni at 53.27, and ranks first on the aggregate score of every evaluation group: image (77.55), video (64.99), visual documents (78.26), text (47.15), audio (50.08) and agent tasks (45.52). Earlier omni-modal embedding systems typically add an audio pathway to a text-vision embedder or align a separately trained audio encoder to an existing embedding space; this work starts directly from a Qwen-Omni understanding model that natively processes text, image, video and audio, without modality-specific projection heads. Evidence comes from the 190 tasks of MMEB-v3, reported as group aggregates and sub-task scores alongside per-entry comparisons with LCO-Embedding-Omni-7B, omni-embed-nemotron-3b, e5-omni-7B and Tianmu-Emb-Uni; across 31 aggregate and sub-task entries the model ranks first on 22 and second on 8.

On the audio- and video-focused suites, Ovis-Embedding-Omni-3B reports MAEB Mean(Task) 57.29 and Mean(Type) 61.22, and MVEB Mean(Task) 61.77 and Mean(Type) 59.72, exceeding the respective runners-up by 3.75/4.10 and 4.19/3.49 points. The work examines fine-grained audio and video discrimination inside one shared embedding space rather than letting acoustic representations adapt to a geometry learned without them. MAEB covers 30 tasks and MVEB 23 tasks, reported as task means and task-type aggregates; the authors note Omni-3B has not yet been submitted to the official leaderboard and that the displayed ranks are estimated by inserting local results into a leaderboard snapshot and recomputing its Borda ranking.

The vision-language variants reach 81.13 overall on MMEB-v2 for Ovis-Embedding-VL-9B and 77.46 for VL-2B, each the best in its comparison group; moving from 2B to 9B raises the overall score by 3.67 points, with the largest gain of 5.78 points on video. The two variants are initialized from Qwen3.5-2B and Qwen3.5-9B and share the same minimal adaptation of removing the language-modeling output head and pooling the final-layer state at the last non-padding token, covering different accuracy and efficiency requirements. The MMEB-v2 overall score is the unweighted average over 78 constituent datasets, with modality-level and sub-task scores reported against seed1.6-embedding-1215, Qwen3-VL-Embedding-8B, DME-Medium and Octen-VL-Embedding-Large.

The training recipe has four stages: low-rank contrastive pretraining, full-parameter homogeneous-source finetuning, annealing embedding distillation, and elastic-dimension adaptation at inference; the authors report that halving the embedding from 2,048 dimensions is essentially free and that 128 dimensions, a sixteen-fold reduction, still retains 93.2% of the average score, versus 85.8% for naive truncation. Homogeneous-source sampling draws each micro-batch from one dataset to obtain task-consistent hard negatives; focal loss reweights queries by current retrieval difficulty; embedding distillation transfers the teacher's full soft distribution over the candidate set; elastic dimensions use a shared orthogonal transform plus a zero-initialized residual linear adapter, avoiding mixing multi-width objectives into encoder training. Elastic-dimension results are presented across the MMEB-Text, Image, Video, Audio, VisDoc and Agent suites at widths 2,048/1,024/512/256/128, with naive truncation as the baseline; the authors note three suites sit marginally above full width at 512 dimensions and read this as no measurable change.

Perspective

The results target retrieval and retrieval-augmented generation settings that need a unified index over text, images, video and audio, including agent use cases such as tool, GUI and knowledge retrieval; the family covers four modalities including audio with Omni-3B and text, image and video with VL-2B and VL-9B, with output dimensions inherited from the backbones at 2,048 and 4,096. Elastic-dimension adaptation targets deployment configurations with different storage, latency and accuracy budgets, and the authors describe the useful range as the shorter embedding prefixes. The authors state they will open-source checkpoints, training and data-construction recipes, inference code and a unified evaluation toolkit, so that audio, audio-video and any-to-any retrieval settings with few existing open resources can be reused directly.

The corpus is described as roughly 50M query-target pairs, but the authors explicitly label the numbers in that subsection as placeholder estimates based on the current data freeze, to be updated for the camera-ready version, so corpus scale and composition should be read against the final version. MAEB and MVEB are used in beta versions, and Omni-3B has not yet been submitted to the official leaderboard, so the displayed ranks come from inserting local results into a leaderboard snapshot. In the elastic-dimension results, three suites sit marginally above full width at 512 dimensions, which the authors read as no measurable change; audio and video lose more at short prefixes than visual documents and agent tasks, which the authors attribute to the information capacity of a low-dimensional vector rather than to the compressor. On MMEB-v3, MultiConIR and memory retrieval are the relatively weaker entries, leaving multi-condition retrieval and long-horizon memory matching as open directions.

Sources