arXiv The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
arXiv SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
arXiv The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
arXiv The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.
The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.
The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.
The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.
arXiv Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
arXiv SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
arXiv The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
arXiv UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
arXiv The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
arXiv The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
arXiv This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
arXiv Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
arXiv The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
arXiv The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
arXiv The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
arXiv The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
Nature News A Nature news story describes a mouse study, published in Advanced Healthcare Materials, in which researchers tweaked the antiparasitic drug niclosamide to target cells implicated in endometriosis, while an npj Women's Health analysis reports that among Denmark's top 100 grant-awarding foundations endometriosis research received EUR 173,958 compared with EUR 254,908,430 for diabetes and EUR 325,940 for inflammatory bowel disease, alongside a wide gap in media mentions.
A Nature news story describes a mouse study, published in Advanced Healthcare Materials, in which researchers tweaked the antiparasitic drug niclosamide to target cells implicated in endometriosis, while an npj Women's Health analysis reports that among Denmark's top 100 grant-awarding foundations endometriosis research received EUR 173,958 compared with EUR 254,908,430 for diabetes and EUR 325,940 for inflammatory bowel disease, alongside a wide gap in media mentions.
A Nature news story describes a mouse study, published in Advanced Healthcare Materials, in which researchers tweaked the antiparasitic drug niclosamide to target cells implicated in endometriosis, while an npj Women's Health analysis reports that among Denmark's top 100 grant-awarding foundations endometriosis research received EUR 173,958 compared with EUR 254,908,430 for diabetes and EUR 325,940 for inflammatory bowel disease, alongside a wide gap in media mentions.
A Nature news story describes a mouse study, published in Advanced Healthcare Materials, in which researchers tweaked the antiparasitic drug niclosamide to target cells implicated in endometriosis, while an npj Women's Health analysis reports that among Denmark's top 100 grant-awarding foundations endometriosis research received EUR 173,958 compared with EUR 254,908,430 for diabetes and EUR 325,940 for inflammatory bowel disease, alongside a wide gap in media mentions.
arXiv The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
arXiv Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
arXiv This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.