Skip to main content
Back to timeline
arXivSource publication:

OmniSeek lets an Omni-LLM decide when to look or listen across turns, consistently improving audio-visual reasoning on multimodal benchmarks

Related research and updates

Synopsis

The authors present OmniSeek, an agentic framework that makes evidence acquisition part of reasoning: the model dynamically decides whether to look or listen and over which temporal window within a multi-turn protocol, appends retrieved raw audio or visual segments back into context, cold-starts this behavior by supervising on the synthesized OmniTraj-170K corpus of multi-hop Chain-of-Thought trajectories, then optimizes the policy via two-stage reinforcement learning with verifiable rewards, and adds an Audio-Visual Necessity objective that rewards successful trajectories whose reasoning depends on both modalities, with experiments reporting adaptive cross-modal evidence seeking and consistent gains in audio-visual reasoning.

Source-provided article image: OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
Figure 1 ·

Figure 1 : Interleaved Multi-turn Audio-Visual Reasoning. OmniSeek acts as an active agent operating through an iterative <think> → \rightarrow <tool_call> → \rightarrow <observe> loop. The agent dynamically alternates between fetching audio cues and extracting visual evidence from different time spans to answer a complex multi-hop question, avoiding the pitfalls of single-modality shortcuts.

arXiv

Interpretation

OmniSeek makes evidence acquisition part of the reasoning process rather than passively processing an entire audio-visual sequence in a single forward pass. In contrast to single-pass processing of full audio-visual sequences, the framework lets the model dynamically decide whether to look or listen and over which temporal window, retrieving sparse but critical cross-modal evidence within long contexts. The abstract describes the framework and multi-turn protocol; it does not name specific benchmarks, report numbers, or present ablations.

Through an iterative multi-turn protocol, retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. Feeding retrieved evidence back into context grounds later reasoning steps in acquired raw evidence, forming a multi-turn loop. Mechanism-level description at the abstract level; no quantitative detail on number of turns, context length, or retrieval granularity.

The authors build a data engine that synthesizes OmniTraj-170K, using multi-hop Chain-of-Thought trajectories to cold-start multi-turn tool-use behavior. Supervision on trajectories with interleaved audio and visual evidence is followed by two-stage reinforcement learning with verifiable rewards to further optimize the policy. The abstract gives the corpus name and scale and the order of the training pipeline, but does not report data composition ratios or training hyperparameters.

An Audio-Visual Necessity objective explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. The reward design adds a requirement for dual-modality dependence rather than relying only on task success as the reward signal. The abstract states the direction of this objective but gives no reward weights, control settings, or quantified reduction in shortcut behavior.

Perspective

The work targets multi-turn audio-visual reasoning where sparse critical evidence must be retrieved from long audio-visual contexts, and it applies to Omni-LLM agents with tool-calling ability. It enables the model to actively choose whether to look or listen and over which temporal window during reasoning, and to append raw segments back into context, offering a reusable training and optimization paradigm for long-context multimodal question answering and cross-modal evidence localization; OmniTraj-170K and the two-stage verifiable-reward reinforcement learning pipeline also provide a reference for cold-starting multi-turn tool use.

Based on the abstract alone, the specific benchmarks, evaluation metrics, magnitude of improvement, ablation results, and the weighting of the Audio-Visual Necessity objective cannot be confirmed; the synthesis method, data quality, and coverage of OmniTraj-170K, as well as the stability and reproducibility details of the two-stage reinforcement learning, still need to be checked in the full text. How multi-turn tool use behaves under longer contexts or higher temporal resolution, and whether residual reliance on a single modality persists, are open questions worth watching.

Sources