Skip to main content

AI Core

974 items

  1. arXiv

    SplitMoE splits the video-diffusion expert pool into semantic and generic branches, beating a same-source MoE at 14B activated parameters and showing coarse-to-fine denoising routing

    The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
  2. arXiv

    TRM first writes a case-adaptive rubric before scoring, beating open-source reward models and nearing proprietary ones on image generation and editing benchmarks

    The work introduces the "Think Before You Score" paradigm and the Thinking Reward Model (TRM), which first generates a case-adaptive rubric for each task condition and candidate output, then inspects the candidate criterion by criterion and aggregates the evidence into a fine-grained pointwise reward; it also proposes PD-GRPO to use pairwise preference supervision for better discrimination while mitigating score polarization, and reports state-of-the-art results among open-source reward models on image generation and editing reward-modeling benchmarks, highly competitive with proprietary alternatives, with TRM-guided reinforcement learning consistently improving diverse visual generation models.
  3. arXiv

    Turning context into an editable file: CLM lets models manage their own context, raising BrowseComp-Plus accuracy by 11.4% with 21.5% fewer FLOPs

    The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
  4. arXiv

    APM-Bench tests cross-session persistent memory for streaming video assistants across 549 sessions and 104 life trajectories, finding existing methods struggle to combine recall, latency, and storage

    The work introduces APM-Bench, which reformulates egocentric streaming interaction as multi-session life trajectories (549 sessions, 104 trajectories, 2,719 candidates, averaging 69 minutes of video per trajectory) and evaluates general video models and eight specialized memory systems on cross-session understanding, real-time perception, and adaptive response, plus an evidence-availability-aware test; results reveal a clear utility-latency-storage trade-off: raw video memory gives the strongest cross-session performance but needs GiB-scale storage and high latency, text summaries cut storage to the KiB scale at the cost of cross-session performance, event-structured memory shows the strongest utility among specialized systems, and adaptive response remains difficult even with rich history
  5. arXiv

    ReImaGin uses image generation as a multimodal reasoning tool, beating text-only reasoning and specialist vision-tool baselines by up to 25% across six visual reasoning tasks

    The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
  6. arXiv

    SEAD recasts tool-agent attack and defense as partially observed state control: DART lifts semantic attack success by 18.8–35.9 points, SAGE cuts executable attack success from 48.0% to 4.0%

    The work formulates attack and defense for tool-using language-model agents as partially observed state control (SEAD), derives from it an execution-feedback tree-search attacker (DART) and a pre-execution defender that can run read-only state investigations (SAGE), and reports across 187 malicious tasks and four target models that DART exceeds baselines by 18.8–35.9 percentage points under both semantic and executable criteria, while SAGE preserves 95.79% of benign trajectories, intercepts 92.73% of harmful paths on recorded trajectories, and reduces DART's executable attack success from 48.0% to 4.0% online.
  7. arXiv

    Adversarial post-training restores missing high-frequency detail in pixel diffusion, cutting DeCo FID from 33.27 to 28.59

    By keeping the original diffusion or flow-matching objective and adding an adversarial loss to the predicted output at non-high-noise timesteps, this study jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality on the DeCo and PixelGen pixel backbones, attributing the gain to restored missing natural-image high-frequency power, while the same procedure yields no comparable gain in the tested latent diffusion configurations.
  8. arXiv

    StoryEngine constrains multi-shot video generation with an explicit story world state, lifting anchor persistence to 0.9389 on a 60-story benchmark

    The work proposes StoryEngine, a state-grounded agentic framework that separates an authoritative semantic plan from fallible generated pixels, using story-state propagation, grounded render planning, and a bounded evaluation-guided repair loop to produce long-form multi-shot video; on a self-built 60-story benchmark with two video backbones, Veo 3.1 and Wan2.2-TI2V-5B, it outperforms ViMax, MovieAgent, and the planner-free Direct I2V baseline on all metrics of narrative realization, cross-shot coherence, and visual consistency.
  9. arXiv

    Looped language models gain reasoning but lose knowledge when unrolled beyond the training horizon, and history-state injection with timestep conditioning mitigates the trade-off

    Through controlled pretraining experiments, this study systematically examines when recurrence helps, where it should be applied, and how it should be conditioned in looped language models (LoopLMs), finding that extra loops beyond the training horizon improve reasoning (e.g., a reasoning score rising from 28.52 to 31.49) while degrading knowledge (from 62.80 to 52.11), that non-recurrent output layers improve robustness to under-unrolling, and that the proposed history-state injection (especially channel-wise) combined with timestep conditioning better preserves knowledge under extended unrolling and improves robustness across inference budgets.
  10. arXiv

    AgentTell benchmark shows browser-use agents leak private user information through click choices in 61.1% of sessions, and falsely assure users of privacy in 34.5% of leaking sessions

    The authors define and formalize behavioural side-channel leakage in browser-use agents, introduce the AgentTell benchmark of 20 scenarios and 100 tasks, and evaluate six backbones across 9,760 sessions, finding that agents carrying a secret reveal it through task-directed action choices in 61.1% of sessions despite an explicit privacy instruction, that agents still leak in 56.7% of sessions where their own memory states the secret must not be shared, and that in 34.5% of leaking sessions their final responses falsely assure users that no information was disclosed.

Page 22 · showing 10