Omni-Embed-Mini binds speech, audio, image, video and visual documents into one cosine space via a frozen text backbone, holding text retrieval at 49.57 nDCG@10 with 0.9B parameters
Synopsis
MBZUAI presents Omni-Embed-Mini, which recasts cross-modal alignment as self-distillation through a shared frozen backbone: each media sample is paired with a dense cascaded caption and the teacher target is simply that same frozen backbone's embedding of the caption, so only projectors and phased LoRA adapters are trained; the 0.9B variant maps text, speech, audio, images, video and visually-rich documents into one cosine space while keeping text weights bit-identical, scoring 49.57 nDCG@10 on MTEB-v2 BEIR-8, and the 2.3B variant edges ahead of the closed gemini-embedding-2 on the overall-modality average, 51.39 to 49.51.
Interpretation
The paper splits the multimodal expansion dilemma into two coupled failure modes, catastrophic forgetting on the text side and modality bloat, and offers a recipe that updates no text-side parameter. Earlier omni embedders (LCO-Embedding-Omni-3B/7B, BidirLM-Omni-2.5B, omni-embed-nemotron-3B, e5-omni-3B/7B) absorb new modalities through joint contrastive training and scale to billions of parameters; here text is an immovable anchor and new modalities enter only through external modality encoders and projectors. The paper reports 935.35M inference parameters for the 0.9B variant, of which 869.98M are frozen and 68.32M trainable (7.3%); the 2.3B variant trains 141.68M (5.8%), with a parameter comparison table against each open omni embedder.
The teacher signal needs no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is the frozen backbone's own EOS-pooled embedding of that caption, so teacher and student share weights and inhabit byte-identical geometry. ImageBind binds through images and LanguageBind through language but with a separate frozen text tower; here teacher and student are the same untrained weights, so no projection head is needed and targets are fully cacheable. The paper states that because the backbone is frozen the teacher embedding is constant across training and is computed once and cached on disk under a stable sample id; the alignment loss is cosine self-distillation in the native pooled hidden space, with full loss equations given.
The training recipe combines a Matryoshka SigLIP pairwise sigmoid contrastive loss, an online hybrid hard-negative miner and phased LoRA, and carries over from 0.9B to 2.3B. The online miner follows ANCE's idea of refreshing the negative pool against the current encoder but generalises it to multiple modalities and Matryoshka: the text index is perfectly stable because the text branch never receives gradient, while only the media-side index drifts; the SigLIP pairwise loss performs well at small batch size. Ablations show only combined text and media mining separates clearly from no mining (+1.81 on the mean over the five media modalities), visual-document retrieval is the least mining-sensitive modality (four configurations agree within 0.6 nDCG@10 on ViDoRe-V3), and across three seeds text is identical at 49.57 while media standard deviations are 0.43 image, 0.59 general audio, 0.95 video, 1.03 speech and 1.76 visual document.
Dense captions are the single most important data-side ingredient, and replacing them with the datasets' original short captions costs substantial points. The paper treats data preparation as a first-class methodological contribution, generating cascaded captions with Qwen3-Omni-30B-A3B-Instruct at low temperature and interpolating the source dataset's ground-truth caption as a grounding prior in the system prompt. Table 2 shows dense captions over original captions gain +2.68 speech, +2.70 audio, +8.60 image and +14.06 video, 30.79 versus 23.78 overall; on visual documents the swap costs 4.8 nDCG@10 (46.98 to 42.21).
Perspective
The result targets multimodal retrieval in a research-prototype setting: text, speech, audio, images, video and visually-rich documents are mapped into one cosine space, and the text inference path touches only unchanged backbone weights, so it suits deployments that want to add modalities without sacrificing existing text retrieval quality. The 0.9B variant fits size-sensitive settings that can accept speech and audio retrieval trailing larger baselines; the 2.3B variant is stronger on video (55.18), image (64.80) and visual documents (58.10), suiting applications that need visual-modality quality. Because the backbone is frozen, future text-embedder improvements can in principle be adopted by re-running only projector and LoRA training, as the swap from the 0.6B text encoder to Qwen3-VL-Embedding-2B illustrates. Matryoshka truncation lets a 10M-item float32 index shrink from 41 GB to 5 GB at 128 dimensions, which matters for large-scale indexing.
Open questions the paper itself flags are worth watching: every cascaded caption is machine-generated, and the authors state they do not filter captions for factuality, do not verify them against source media, and did not run a controlled caption-perturbation study, so how far retrieval degrades as caption quality falls cannot be quantified; the only automatic check is a 0.92 cosine threshold used to discard likely false negatives during mining. The audio pathway uses a fixed temporal interleaving of speech and environmental-audio tokens, which may be suboptimal for inputs dominated by one audio type. Visual-document dense captions average about 4,157 characters versus about 582 for audio, and the backbone's 2,048-token context truncates the long tail; truncation is consistent between teacher and student paths so cosine alignment is unaffected, but the upper end of visual-doc detail is lost. On MMEB-V2 all omni-style models embed the benchmark's own 8 frames as images and mean-pool, which keeps comparisons uniform but caps achievable video performance. The 0.9B and 2.3B variants use different vision strategies (a ViT extracted from Qwen3.5-0.8B versus the backbone's native vision module), and cross-variant ablations of the two are left to future work. Beyond that, existing benchmarks each probe a single modality slice and none evaluates any-to-any retrieval across all six modalities simultaneously, so results may not generalise to all languages, domains and real-world retrieval workflows.
