Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OmniSeek lets a 30B omni-modal model decide whether to look or listen: leading same-scale open models across 10 audio-visual benchmarks, 74.6% on VideoHolmes

The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
arXiv

OmniSeek lets an Omni-LLM decide when to look or listen across turns, consistently improving audio-visual reasoning on multimodal benchmarks

The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.