AI Core
970 items
OmniTaskonomy maps 19 generation tasks against 25 understanding capabilities: I2I-then-I2T training improves understanding, and gradient alignment tracks the gains
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
CaptchaArena trains a single CaptchaAgent policy on 20,000 execution-verified CAPTCHA puzzles, lifting average Pass@1 from 11.4 to 71.7 against a human 94.1
The work builds CaptchaArena, a large-scale fine-grained computer-use training dataset of 20,000 interactive CAPTCHA puzzles across 20 types and five interaction modes, where every solution is replayed in a real browser and accepted by the page's own verifier, together with 20,000 screenshot-action trajectories (18,000 carrying judge-filtered step-by-step reasoning annotations) and pixel-mask supervision for irregular targets; training a single 9B policy, CaptchaAgent, on it reaches 70.5 average Pass@1 after supervised fine-tuning and 71.7 after reinforcement learning with the environment verifier as reward, versus 11.4 for the untrained backbone, 35.2 for the strongest open-weight GUI agent, 69.2 for the strongest closed-source model, and 94.1 for humans.
REST adds four differentiable losses to latent thoughts, lifting accuracy by up to 7.5 points across 7 benchmarks
The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
KAIST team turns prompt-template disagreement into preference supervision, lifting open-vocabulary segmentation across the MESS benchmark without pixel-level labels
The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
Ant Group proposes Marathoner: synthesizing tasks from million-line-scale GitHub PRs lets a 9B open model work 10+ hours and make 1000+ tool calls
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
SplitMoE splits the video-diffusion expert pool into semantic and generic branches, beating a same-source MoE at 14B activated parameters and showing coarse-to-fine denoising routing
The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
TRM first writes a case-adaptive rubric before scoring, beating open-source reward models and nearing proprietary ones on image generation and editing benchmarks
The work introduces the "Think Before You Score" paradigm and the Thinking Reward Model (TRM), which first generates a case-adaptive rubric for each task condition and candidate output, then inspects the candidate criterion by criterion and aggregates the evidence into a fine-grained pointwise reward; it also proposes PD-GRPO to use pairwise preference supervision for better discrimination while mitigating score polarization, and reports state-of-the-art results among open-source reward models on image generation and editing reward-modeling benchmarks, highly competitive with proprietary alternatives, with TRM-guided reinforcement learning consistently improving diverse visual generation models.
Turning context into an editable file: CLM lets models manage their own context, raising BrowseComp-Plus accuracy by 11.4% with 21.5% fewer FLOPs
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
APM-Bench tests cross-session persistent memory for streaming video assistants across 549 sessions and 104 life trajectories, finding existing methods struggle to combine recall, latency, and storage
The work introduces APM-Bench, which reformulates egocentric streaming interaction as multi-session life trajectories (549 sessions, 104 trajectories, 2,719 candidates, averaging 69 minutes of video per trajectory) and evaluates general video models and eight specialized memory systems on cross-session understanding, real-time perception, and adaptive response, plus an evidence-availability-aware test; results reveal a clear utility-latency-storage trade-off: raw video memory gives the strongest cross-session performance but needs GiB-scale storage and high latency, text summaries cut storage to the KiB scale at the cost of cross-session performance, event-structured memory shows the strongest utility among specialized systems, and adaptive response remains difficult even with rich history
Page 21 · showing 10