Omni-Decision replaces dialogue history with an evidence ledger, reaching 81.4% on OmniGAIA at about 43% of Gemini-3.1-Pro's cost
Synopsis
The work presents Omni-Decision, an omni-modal agent that replaces the growing dialogue history with an explicit evidence ledger, where a critic reads each noisy observation and passes only usable content to the ledger so the planner works from a compact context throughout the task; each run records state, action, and verdict at every step, and supervised fine-tuning plus decision-level reinforcement learning on these trajectories further improve the planner, yielding state-of-the-art 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's cost per question and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.
Figure 1: Omni-Decision on an OmniGAIA task. Each tool call is digested into ledger updates that close open evidence needs; once all needs are closed, the system answers against the recorded evidence.
arXivInterpretation
The paper locates the main bottleneck of omni-modal agents in planning rather than perception: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Rather than focusing on perception or end-to-end models, it offers a diagnosis through controlled backend replacements, where replacing the planner causes a much larger performance loss than replacing the perception backend. Evidence comes from the controlled backend-replacement comparison described in the abstract; the abstract does not give the specific replacement configurations, sample sizes, or loss magnitudes.
Omni-Decision replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict, while a critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest. This keeps the planner working from a compact context throughout the task instead of letting raw multimodal observations pile up in history. Evidence is the system design and mechanism description in the abstract; the abstract reports no ledger size, critic filtering ratio, or ablation data.
Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Using decision-level trajectories as training signal makes the planner itself improvable, rather than relying only on prompting or a fixed policy. Evidence is the training procedure described in the abstract; the abstract gives no training-data scale, hyperparameters, or before-and-after numbers.
Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model. Compared with prior work, the result positions the system on both accuracy and cost, and matches end-to-end models on long-video understanding. Evidence is the two benchmark scores and the cost ratio reported in the abstract; the abstract gives no confidence intervals, repetition counts, or per-benchmark evaluation details.
Perspective
The result targets multi-step question answering that requires gathering evidence across video, audio, web pages, and computation, and applies to omni-modal agent deployments seeking high accuracy at lower cost per question; the evidence-ledger and critic-filtering mechanisms can also inform the design of other long-horizon multimodal tasks.
The abstract does not give the specific configurations and loss magnitudes of the controlled backend replacements, the ledger size and critic filtering ratio, the training-data scale and hyperparameters, or confidence intervals and repetition counts; moreover, the loaded text is the abstract and metadata, missing the body figures and ablation results, so these mechanism details and the stability of the reported scores should be checked against the original.
