arXiv Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
arXiv The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
arXiv Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
arXiv The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
arXiv The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
arXiv Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
arXiv The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.
arXiv The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
The work shows that training latent recursive LLM systems with only final-answer cross-entropy leaves the thought representation subject to four failures, and introduces REST, which turns causality, minimality, separability and stability into differentiable losses added to cross-entropy, improving accuracy over the CE-only baseline by an average of 3.5 points in the multi-agent setting and 3.3 in the single-agent setting, up to 7.5 and 6.5 points at best, and raising the rate of convergence on a final answer from 73% to 95%, across 7 benchmarks in mathematics, science, medicine and code generation under matched training data and latent budget.
arXiv The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
The work formalizes a failure mode it calls semantic thrashing, in which append-only working memory keeps accumulating noise and dilutes key evidence, and proposes VideoLoop: an outer multimodal agent explores the video inside a sandboxed filesystem while an inner memory orchestrator retrieves relevant artifacts from an unbounded filesystem and rewrites a bounded working memory after every step; across VideoMME (long), VideoMMMU, and LongVideoBench (long), VideoLoop improves four LVLM backbones in a plug-and-play manner, averaging a 4.2-point gain over baseline on VideoMME (long) and reaching 88.3%, 88.8%, and 80.9% with Gemini 3.1 Pro.
arXiv The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
The work introduces AsyncLLM, a general framework built on the Python asyncio async/await programming model that lets pretrained LLMs observe, reason and act as concurrent coroutines, sharing memory in real time through CacheBlocks (attention KV caches and Gated DeltaNet recurrent states) and cache views, and demonstrates asynchronous operation with the same family of Qwen 3.x vision-language models on streaming video understanding, ViZDoom games and DevOps-Gym system monitoring without task-specific training.
arXiv The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
The work identifies two critic failure modes that destabilize PPO for LLM reinforcement learning — filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, and heterogeneous return noise across prompts lets high-variance prompts dominate critic updates in finite batches — and introduces EasyPPO, which combines actor-only overlong filtering, noise-normalized critic regression weighting each prompt by the inverse standard deviation of its sampled returns, and moderately smaller critic mini-batches with gradient clipping, remaining stable across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24 and multi-turn search on Search-R1, with best validation scores improving over PPO by 14.89%, 2.
arXiv The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
The work proposes Scaffolding Minds, a two-stage framework in which Stage 1 replaces the frozen off-the-shelf vision encoder with a scaffolding encoder trained end-to-end through the downstream task loss to produce the latent target, and Stage 2 replaces deterministic latent regularization with a Gaussian sampler whose mean and variance are learned so that residual latent actions can be sampled for RL exploration, improving over the strongest latent-reasoning baseline by 9.5 points on average on FrozenLake (19 points at Level 32) and by 5.6 points on average across nine visual-centric reasoning benchmarks.
arXiv The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
The work identifies a "uniformity trap" in which language-style token-wise MoE routing, combined with uniform expert-usage regularization, scatters neighboring patches of the same object across different experts in video diffusion, causing routing fragmentation and structural distortion; it proposes SplitMoE, which bifurcates the expert pool into semantic and generic experts and uses prototype-guided routing with pull-push regularization so tokens cluster by semantic attributes, outperforming traditional load-balanced MoE on VBench-2.0 and T2V-CompBench under an equivalent activated-parameter budget, while revealing a coarse-to-fine denoising pattern in which semantic-expert weight peaks in the early high-noise stage around steps 10-15 and then shifts toward generic experts.
arXiv The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
The work proposes WorldAttention, an attention architecture for interactive video world models that combines a Hierarchical KV Cache (HKV) storing and retrieving page-level KV across GPU, CPU, and NVMe tiers with a Hybrid Sparse Attention (HSA) that fuses a linear global branch and a head-adaptive block-sparse branch, reporting subject consistency of 0.9472 on VBench-Long and 0.9668 on InterVBench, a 14.02x sparse-kernel speedup over FlashAttention-3, a 2.21x end-to-end speedup, and 22.0 FPS on a single NVIDIA H100.
arXiv The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.
The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.
The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.
The work introduces heterogeneous refinement, giving different feature groups distinct refinement budgets across depth in pixel-space diffusion Transformers, which spontaneously yields persistent features that encode global structure and active features that encode high-frequency detail; building on this, Persistence Forcing (PerF) and Persistence Guidance (PG) reduce FID on ImageNet 256 from 3.66, 2.36, and 1.86 to 2.81, 1.91, and 1.63 for JiT-B/L/H, and on ImageNet 512 from 1.94 to 1.76 for JiT-H/32.
arXiv The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
Nature News Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
Ibáñez and colleagues report in Science Advances a 'speech clock': they recorded 2,928 Spanish speakers from Argentina, Chile, Colombia, Mexico and Peru across various speech tasks, used machine learning to extract more than 700 speech features that change with ageing and dementia (such as pitch and vocabulary range), trained a model to predict age and compute a 'speech age gap', and found the clock could distinguish healthy individuals from those with some form of cognitive impairment, rating the speech of people with cognitive issues as older than expected for their chronological age while healthy people's speech generally matched their age.
arXiv The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
The work introduces the Trajectory Adaptive Progress–Fluctuation Scheduler (TAPS), which uses exponential moving averages to estimate persistent progress and centered fluctuation in recurrent updates and adapts the per-step update scale online; it proves that under stated conditions TAPS reduces expected terminal loss and reaches a target quality in fewer loops, and empirically improves terminal accuracy with up to roughly 1.5x wall-clock speedup across Sudoku, Maze, language-model recurrence, and intermediate-layer recurrence.
arXiv WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
WorldLine introduces an action-driven visual simulator that decouples transferable robot-object dynamics learning from heterogeneous action grounding: it pretrains on more than 10,000 hours of action-free robot videos, grounds them with over 2,000 hours of action trajectories across more than ten embodiments through image-space action maps, adds multi-view, failure-enriched and relational-regularized training, and distills a robot-focused few-step causal rollout; across held-out and out-of-domain settings it maintains visual quality and robot-motion agreement, improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, predicts trajectory success with 74% mean accuracy on RoboTwin and AgiBot, and improves task success by up to 21.
arXiv Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.
Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.
Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.
Holding bytes identical and changing only whether forged chat-template markers are encoded as reserved token ids or as ordinary subwords, this study measures the resulting gap in indirect prompt-injection success on LLM agents in InjecAgent, finding an identity gap of 39 to 66 percentage points on Llama-3.1, GLM-4.5 and Seed-OSS-36B while Qwen3-8B draws most of its attack from marker text, and further showing that the authority lives in the single learned input vector at the marker position and that the standard tokenizer-side mitigation misses the tool-protocol tokens that carry tool output.