OmniSeek lets a 30B omni-modal model decide whether to look or listen: leading same-scale open models across 10 audio-visual benchmarks, 74.6% on VideoHolmes
Synopsis
The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
Interpretation
OmniSeek reframes audio-visual reasoning from single-pass passive encoding into multi-turn active evidence seeking: inside a <think> <tool_call> <observe> loop the model itself decides whether to inspect audio or video and over which temporal window, and retrieved raw segments are appended directly back into the context. Earlier multimodal chain-of-thought methods such as OmniVideo-R1 and OmniReasoner rely mainly on textual reasoning traces, coupled modalities, or single-turn retrieval; OmniSeek uses two decoupled tools for independent cross-modal and cross-temporal routing. The paper defines get_audio_clip(start, end) and get_video_clip(start, end, fps, resolution) and states that loop termination comes from model self-reflection rather than hardcoding; against a text-only CoT baseline given the same two-stage RL training, the multi-turn multimodal paradigm gains +13.5% on OmniVideoTest, +8.1% on WorldSense, and +7.0% on Daily-Omni.
The authors build OmniTraj-170K: 169,725 trajectories over 39,797 videos, where each question is tied to 2 to 7 timestamped evidence spans and must include at least one audio span and one video span. The data engine runs in three stages (structured audio-visual alignment, evidence-grounded QA generation, interleaved trajectory assembly); tool calls and observations are deterministically constructed from the evidence chain while the model generates only the <think> nodes, reducing annotation hallucination. Statistics show 76.1% of trajectories need two tool calls and 23.9% need three or more; most evidence spans last 3 to 10 seconds; source videos come from FineVideo across 122 categories. In a controlled single-turn SFT comparison, training on a 100K subset gains +5.8% on LVOmni and +3.9% on Daily-Omni over the base model.
The three-phase pipeline (cold-start SFT, GSPO reinforcement learning, hard-example refinement) first degrades then surpasses the base model: Phase 1 shows an alignment tax with drops on most benchmarks, Phase 2 RL recovers strongly, and Phase 3 with the Audio-Visual Necessity reward adds further gains. Audio-Visual Necessity runs two no-grad teacher-forced passes with attention masking to measure per-token log-likelihood drops per modality, then uses a logical AND gate to credit the weaker modality, suppressing single-modality shortcuts without extra rollouts. Ablations show Phase 1 lowers Daily-Omni from 71.9% to 69.1% and WorldSense from 55.1% to 50.3%; after Phase 2 VideoHolmes rises from 55.9% to 69.6%; adding the necessity reward over the standard RL baseline yields a further +3.2% on WorldSense and +2.2% on OmniVideoTest.
Across 10 omni-modal benchmarks OmniSeek reaches leading or competitive open-source performance, with the largest margins on long-form and sparsely cued scenarios. Relative to the base Qwen3-Omni-Instruct it gains +16.3% on MMOU and +8.4% on LVOmni; relative to the text-reasoning model OmniVideo-R1 it scores 74.6% vs 62.9% on VideoHolmes and 47.7% vs 44.8% on OmniVideoBench. Numbers come from the paper's tables: Daily-Omni 80.0, AVUT 78.8, WorldSense 62.4, FutureOmni 58.3, OmniVideoTest 69.5, VideoHolmes 74.6, JointAV 72.8, OmniVideoBench 47.7, MMOU 70.4, LVOmni 44.2; general video benchmarks Video-MME 78.5, MLVU 77.1, LongVideoBench 66.4, LVBench 51.4.
Perspective
The work targets long audio-visual question answering that requires cross-modal, cross-temporal evidence seeking, and applies to omni-modal large models with tool-calling ability; its data engine and training pipeline can be reused by other Omni-LLMs, and OmniTraj-170K plus the Audio-Visual Necessity reward give later work a directly usable starting point.
The paper itself reports a structural bottleneck on extremely long videos (e.g., 2 hours): audio tokens scale linearly with time, and the base model's 32K context is exceeded by roughly 90K audio tokens for two hours of audio, after which the model falls into repetitive <think> loops and fabricates observations instead of calling tools; the proposed linear speed-up heuristic may distort pitch and environmental sound texture, and native context extension plus modality-asymmetric token compression remain future work. In addition, on OmniVideoTest the OmniVideo-100K training data has an advantage because it shares origin and distribution with the benchmark, so cross-distribution generalization still needs more independent benchmarks.
