Skip to main content

Search

“All disciplines” · 1436 results

Page 17 · showing 20
arXiv

OmniTaskonomy maps 19 generation tasks against 25 understanding capabilities: I2I-then-I2T training improves understanding, and gradient alignment tracks the gains

Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
arXiv

Shared or dual projections? A bias–variance boundary gives the criterion, and CARS cuts held-out regret by 49–96% across five datasets

The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
arXiv

Raven composes model–harness pairs into a multi-agent ecosystem and leads planning, four specialist domains, and skill reuse over Claude Code and other baselines

Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
arXiv

ByteDance team finds chunked KV-cache compression makes long-context retrieval periodically weak at the compression stride, with up to 40 percentage points between phases in DeepSeek-V4

The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
arXiv

Ant Group proposes Marathoner: synthesizing tasks from million-line-scale GitHub PRs lets a 9B open model work 10+ hours and make 1000+ tool calls

The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
arXiv

Ranking positions by chat structure lets auditors explain 5% of tokens and keep nearly all threat-detection success

Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
arXiv

AutoDataBench makes agents deliver training tasks one at a time: five frontier agents all score below 20 out of 100 at 45 minutes, with difficulty calibration rather than mode coverage as the binding constraint

The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
arXiv

REST adds four differentiable losses to latent thoughts, lifting accuracy by up to 7.5 points across 7 benchmarks

The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
arXiv

VideoLoop uses dual-loop bounded working memory to ease semantic thrashing in long-video agents, lifting Gemini 3.1 Pro by 3.2 to 4.5 points on VideoMME (long) and two other benchmarks

The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
arXiv

Yandex team's AsyncLLM lets Qwen 3.x models watch, think and act concurrently without fine-tuning, speeding up streaming video, games and system monitoring

The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
arXiv

EasyPPO traces two critic failure modes in PPO and stays stable across three tasks, gaining 14.89%, 2.28% and 9.47% over PPO

The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
arXiv

Scaffolding Minds swaps in a learnable scaffolding encoder and an adaptive Gaussian sampler for latent visual reasoning, gaining 9.5 points on FrozenLake and 5.6 points across nine visual reasoning benchmarks

The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
arXiv

SplitMoE splits the video-diffusion expert pool into semantic and generic branches, beating a same-source MoE at 14B activated parameters and showing coarse-to-fine denoising routing

The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
arXiv

WorldAttention pairs hierarchical KV caching with hybrid sparse attention to reach 22 FPS on a single H100 and 0.9472 subject consistency on VBench-Long

The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
arXiv

PerF makes pixel-space diffusion Transformers specialize by feature: FID on ImageNet 256 drops from 1.86 to 1.63 with about 3.6% more parameters

The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.
arXiv

Opus 5.5's agent-designed libraries cut downstream code and beat the human production library by 2.3 points, yet 11 of 15 tasks merely reproduce human abstractions

The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
Nature News

AI 'speech clock' estimates ageing speed from four minutes of speech and rates cognitively impaired voices as older

Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
arXiv

TAPS adapts recurrent step size to the trajectory, lifting Sudoku accuracy to 91.39% and Maze to 79.90% while cutting the loops needed for matched quality

The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
arXiv

WorldLine pretrains on 10,000 hours of action-free robot video and grounds 2,000 hours of action trajectories across ten embodiments, lifting robot-mask IoU by 0.1626 on failed trajectories and predicting trajectory success at 74% mean accuracy

WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
arXiv

Same bytes, two kinds of authority: splitting forged chat-template markers into ordinary subwords cuts injection success by 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B

Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.