Public articles linked to the same research event.
arXiv The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
The authors present OmniSeek, which turns an Omni-LLM into an active multi-turn tool-using agent (get_audio_clip, get_video_clip), backed by a data engine that synthesizes OmniTraj-170K with 169,725 interleaved audio-visual chain-of-thought trajectories, plus a three-phase training pipeline and an Audio-Visual Necessity reward, achieving leading or competitive results on 10 omni-modal and 4 general video benchmarks.
arXiv The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.
The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.
The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.
The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.