AI Core
993 items
Omni-Decision turns the planning bottleneck of omni-modal agents into an attributable object, reaching 81.4% on OmniGAIA at about 43% of Gemini-3.1-Pro's cost
The work presents Omni-Decision, which replaces the growing dialogue history with a task-scoped evidence ledger: a critic digests each noisy multimodal observation into typed evidence events committed by a deterministic reducer, so the planner always decides over a compact, verified state; controlled backend replacements show that replacing the planner costs far more than replacing the perception backend, and the system reaches 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's per-question cost and 65.0% on WorldSense long-video understanding, with execution trajectories used to fine-tune and reinforce weaker planners.
VideoLoop uses dual-loop bounded working memory to ease semantic thrashing in long-video agents, lifting Gemini 3.1 Pro by 3.2 to 4.5 points on VideoMME (long) and two other benchmarks
The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
One token-level interaction decomposition links data selection, collision-versus-erosion forgetting, and plasticity loss into a single learning-dynamics account
The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
Q&D Trains an 8B Questioner on the Consequences of Its Questions: Required-Evidence Coverage on MuSiQue Rises from 78% to 90%, and Retail Task Success from 13% to 34%
The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
Braco reframes visual-token compression from picking tokens to re-parameterizing them, reaching 95.2% normalized accuracy at 23–64x compression while cutting prefill FLOPs by 84.2%–86.7%
The work recasts extreme visual-token compression as a token-parameterization problem, separating basis transformation and structured truncation (which fix the retained subspace and compressibility) from coordinate organization (which affects optimization and cross-modal alignment, i.e. learnability), and designs Braco, a lightweight four-step coder combining transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling; experiments show Braco forms the favorable empirical accuracy–efficiency frontier under 23x–64x compression and remains competitive at 144x, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%–86.
FocusVTC renders long text as low-DPI pages and zooms only the regions reasoning needs, scoring 87.4 on RULER v1 at roughly 2.9x compression versus 57.5 for Glyph
FocusVTC introduces an adaptive-resolution visual text compression framework that keeps global coverage with low-DPI page images and uses tool calls to re-read reasoning-relevant regions at aligned 144 DPI, trained with multi-resolution supervised fine-tuning and GRPO on 29.4K Reasoning-Evidence Localization chain-of-thought examples that link reasoning traces to page indices and bounding boxes, reaching 87.4 on RULER v1 at 72 DPI (versus 57.5 for Glyph), 56.40 on LongBench (versus 55.86 for its text-input backbone), a 13.91-point MRCR macro-average gain, and a 51.19 VTCBench macro-average, while MMMU rises from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
PMOPD projects parameter updates away from protected task subspaces to ease the multi-teacher distillation capability seesaw, lifting the three-task average by 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B
The work proposes PMOPD, which builds per-matrix low-dimensional subspace memories from the cumulative parameter displacements of completed task blocks and projects both gradients and Adafactor optimizer updates of later tasks to remove components conflicting with protected task directions, while a lightweight conflict probe sets task order and a subspace-consistency criterion sets the cycle count; across Code, Reason, and Math, PMOPD improves every evaluated capability over MOPD, raising the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.
AutoRef lets a coding agent rewrite the harness automatically, lifting frozen FLUX.2 [klein] 4B from 5.72 to 7.37 on four-reference MultiBanana and matching Nano Banana Pro and GPT-Image-1.5
AutoRef has a coding agent iteratively rewrite harness code while keeping both the image generator and the reasoning model frozen, separating the tasks whose feedback informs proposals from those used to select candidates and continuing the search from a beam of top-ranked harnesses; the resulting AutoRef-Harness raises open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 (out of 10) on the held-out four-reference MultiBanana test split, matching or exceeding proprietary models such as Nano Banana Pro and GPT-Image-1.5, and transfers without re-optimization across reference counts, benchmarks and evaluators, generators, and reasoning models.
Tsinghua's Leap Lab finds in a matched comparison that world action models generalize from a single inference-time forward pass, not from denoising a clean future
Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
Page 28 · showing 10