Omni-Decision turns the planning bottleneck of omni-modal agents into an attributable object, reaching 81.4% on OmniGAIA at about 43% of Gemini-3.1-Pro's cost
Synopsis
The work presents Omni-Decision, which replaces the growing dialogue history with a task-scoped evidence ledger: a critic digests each noisy multimodal observation into typed evidence events committed by a deterministic reducer, so the planner always decides over a compact, verified state; controlled backend replacements show that replacing the planner costs far more than replacing the perception backend, and the system reaches 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's per-question cost and 65.0% on WorldSense long-video understanding, with execution trajectories used to fine-tune and reinforce weaker planners.
Interpretation
Controlled backend replacements localize the dominant bottleneck of omni-modal agents on the planning side: with Gemini-3.1-Pro perception fixed, moving the planner from GPT-5.2 through GLM-5.2, Qwen3.5-Plus, and Qwen3.5-27B lowers overall accuracy step by step from 81.39% to 46.94%, and Qwen3-Omni-30B leaves only 14.72%; with the GPT-5.2 planner fixed, all three perception replacements stay at or above 58.89%. Earlier judgments that planning is the shortfall came largely from indirect cross-benchmark observations, such as coding agents without native audio-video perception performing comparably to native models and harness differences for the same model exceeding model differences; this work turns that judgment into an attribution within the same omni-modal task domain via one-backend-at-a-time replacement. Same evaluation window, same judge, everything else held fixed, one backend replaced at a time; the Qwen3.5 Plus tier gives the most direct comparison within one generation, where replacing the planner costs 23.3 points and replacing perception costs 15.8 points, a ratio of approximately 1.5; once the planner is weak enough (replacing both gives 13.89%), perception quality barely affects the outcome.
The evidence-ledger mechanism keeps observation noise out of persistent context: each task maintains a typed ledger of open needs, fact and computation slots, confirmed evidence atoms, and unresolved conflicts; a critic digests each observation into a typed evidence event in a transient call; the reducer is the sole writer and commits deterministic field-by-field updates; and the planner's persistent context contains only the ledger's control projection plus the single raw observation currently under review. Existing systems carry intermediate observations in free-form traces, file workspaces, or untyped textual memory, and planning operates directly on those sequences; compression methods such as ACON and the complexity-trap study compress messages after they enter history, whereas this design digests observations into typed state updates on arrival, keeping the persistent portion approximately constant in length across steps. In the state ablation on OmniGAIA, the full method reaches 81.39%, the rolling natural-language-summary Memory arm 68.33%, and no-ledger ReAct 60.28%; relative to no-ledger ReAct the gain is about 24 points on Medium and Hard versus about 16 points on Easy, indicating that longer tasks depend more on tracking which needs are closed.
Execution trajectories themselves form a training signal that requires no step-level manual annotation: state-SFT retains runs the judge marks correct and expands each step into a ledger-state and planner-action pair, while closure-aligned reinforcement learning targets the dominant failure of weak planners, abstaining or converging while needs remain open, scoring a single decision with rule-based closure checks and a lightweight reward judge without any system rollout. The recipe uses the ledger's explicit list of open needs as the basis for offline scoring of individual decisions, so decision-level RL obtains dense feedback without executing the action; compared with DA-GRPO in Orchestra-o1, two of the four reward dimensions are determined directly by ledger rules and the other two are anchored on the open-need list, reducing the reward judge's discretion. Qwen3.5-27B improves from 46.94% to 50.56% with state-SFT followed by RL, with consistent gains on Medium and Hard; Qwen3-Omni-30B improves from 14.72% to 18.61% with state-SFT alone; the training set contains 640 trajectories (100 OmniGAIA Easy and 540 WorldSense), and the Medium and Hard evaluation sets are disjoint from training data by task ID.
On open-world tool interaction and conventional long-video understanding the system reaches 81.39% and 65.0% respectively: on OmniGAIA it exceeds Gemini-3.1-Pro's 79.44% measured in the same window with the same judge, at about $1.2 versus about $2.8 per question; on WorldSense it is on par with the leading end-to-end models on the public leaderboard (Gemini-3.1-Pro at 65.5%), at about $0.3 versus about $0.8 per question. With the same pair of backends, overall accuracy ranges from 5.56% to 81.39% across harnesses, so the harness decides how much of the models' capability is realized; this work makes that layer an explicit, definable, and ablatable object and reports cost from billing records rather than inferring it from token counts. OmniGAIA contains 360 open-world evidence-seeking questions spanning video, audio, images, web evidence, and computation; WorldSense is a multiple-choice long-video benchmark whose answers are contained in video, audio, and subtitles with web search disabled; costs come from billing records, ranging from $0.46 to $2.77 per question on OmniGAIA and $0.08 to $0.40 on WorldSense.
Perspective
The results target omni-modal question answering that requires seeking evidence across video, audio, web pages, and computation, as well as long-video understanding where answers are contained in media; the ledger mechanism requires no training at inference time, so it can be applied to different planner backbones and supports both native multimodal input and on-demand toolized perception. For practitioners this means the system layer can be defined, ablated, and backend-replaced as an independent object, and the ledger states, actions, and verdicts recorded during execution serve as a post-training signal without step-level manual annotation; for researchers it offers an experimental paradigm that attributes bottlenecks by replacing the planner and perception backends one at a time within the same task domain. The authors also note one scope boundary of the current design: the critic reads observations only in text form and media never enter its input, so ledger quality is bounded by the text output of the perception tools, and a critic that reads multimodal input directly is the natural next step.
Several open questions remain for a careful reader: how the scope boundary that ledger quality is bounded by the text output of perception tools will shift once a stronger multimodal critic exists; whether the training gradient that varies with the planner backbone's initial capability still holds for stronger backbones; how improvements in the remaining failure sources interact with planning-layer improvements, given that the evidence-progress audit concentrates residual failures in media perception and external retrieval (of 67 incorrect cases, 27 media perception, 22 external-fact retrieval, 9 verification or conflict resolution, 6 computation, and 3 premature stopping); and that in the same-window comparison Gemini-3.1-Pro uses built-in retrieval to which the same domain block cannot be applied, so the authors report its measured score without adjustment. This summary is based on the loaded full text and external story and does not include every detail of the figures and appendices.
