Skip to main content

Search

“All disciplines” · 1367 results

Page 6 · showing 20
arXiv

First joint scaling law for looped mixture of experts: sparsity raises and delays the saturation of recurrence's effective-parameter gain

The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
arXiv

SkillGym turns 184k community skills into verifiable training environments, and its 9B fine-tuned model beats a 397B untrained model on two benchmarks

SkillGym proposes an automatic pipeline that crawls skills from the internet and keeps those whose workflows run reproducibly offline, uses a builder-reviewer agent system to construct skill-critical tasks across four reasoning structures each with a reference solution and an executable verifier, then collects successful trajectories from three teacher models and four agent harnesses for supervised finetuning, improving six models from 2B to 122B parameters on 22 of 24 comparisons across four skill-use benchmarks, with the 9B model outperforming Qwen3.5-397B-A17B on the SkillGym test set and SkillEval and raising the rate of reading the relevant skill from 28% to 96%.
arXiv

When one observation admits several valid actions, CVAE KL regularization and flow-model Lipschitz smoothness decide whether a policy keeps multimodality, while standard robot simulation benchmarks turn out to be nearly unimodal

The work formalizes multimodality in behavioral cloning, proves that latent-variable policies preserve demonstrated modes only if the latent carries action-conditioned information and that excessive posterior-prior regularization suppresses it, shows that action-space generative policies are limited by the Lipschitz constant of the base-to-action map, validates these mechanisms on synthetic multimodal navigation and a real-robot bimodal tissue-grasping task, and uses a GMM modal-clustering diagnostic to find limited conditional multimodality in Push-T, UR3, LIBERO, and Meta-World, where deterministic regression remains competitive.
arXiv

AIM treats research ideas as explicit search objects: 67.0% and 55.8% average scores across 10 AutoLab tasks, matching the strongest baseline up to 3.1x faster in wall-clock time

The work introduces the Agentic Idea Manager (AIM), a fully autonomous framework that makes research ideas explicit search objects: a Bayesian-optimization-inspired Agentic Surrogate organizes ideas into semantic clusters and produces ordinal promisingness estimates, an Agentic Acquisition mechanism balances exploration and exploitation at cluster and idea level, a Solution Auditor checks idea-solution integrity, and a Resource Planner adaptively allocates parallel branches under a fixed budget; on 10 AutoLab tasks AIM reaches 67.0% average on System Optimization and 55.8% on long-horizon Model Development & CUDA, exceeding the strongest baseline ScientistOne by 1.6 and 4.9 percentage points and reaching that baseline's best score up to 3.1x faster in wall-clock time.
arXiv

PivotOPD concentrates distillation on pivotal mistakes and the turns after them, lifting Qwen3-1.7B by 5.5 points over the strongest baseline on ALFWorld

Analyzing failed ALFWorld rollouts of three Qwen3 models (8B–235B), the work finds that more than half of failures contain an early and usually recoverable pivotal mistake, and proposes PivotOPD: a teacher model locates pivotal turns and names a gold action plus recovery actions at the next few turns, preventive distillation uses reverse KL on the committed mistake while recovery distillation uses forward KL on self-teacher responses, both folded into a single PPO update with group-based RL, achieving the best average performance on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students and transferring to a Nemotron-3.5 student on SWE-Bench Verified (+3.2% resolve rate).
arXiv

SlideDP speeds shared-host multi-GPU full-parameter fine-tuning by 1.46–2.64x and beats FSDP2's measured peak by 11.2% at a larger batch on four H100s

SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
arXiv

Mid-Harness verifies candidate actions between model and harness, lifting TMAX-9B Pass@1 on TerminalBench-Lite from 50.00% to 68.03%

The work introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before forwarding one for execution while keeping the generator and harness unchanged, and studies action-level test-time compute scaling: on TerminalBench-Lite a GPT-5.6 Sol verifier raises TMAX-9B Pass@1 from 50.00% to 68.03% with 8 sampled actions, wider sampling yields little benefit under weak verification, pairwise verification performs best among the evaluated self-verification mechanisms, distilling the stronger verifier's responses into TMAX-9B further improves Pass@1, and combining action scaling with trajectory scaling (Best-of-N and Sequential Refine) reaches higher success at lower estimated token cost.
arXiv

UniEvo-VL lets a multimodal model teach itself with its own critique, lifting GenEval from 0.747 to 0.808

UniEvo-VL introduces an on-policy self-distillation (OPSD) recipe in which one multimodal model acts as both a student seeing only the vanilla prompt and a teacher conditioned on a critique-derived revised prompt, minimizing per-state divergence between their denoising distributions along the student's own trajectories; on Qwen-Image-2512 it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53, while stronger external critics such as GPT5.6-Luna indicate a higher self-evolving ceiling and text-rendering gains remain uneven.
arXiv

GGSD turns five buttons into playable skills via 1v1 self-play: humans clear Maze and CubePush on Ant, Franka and G1 with no extra training

The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
arXiv

LoopVL reuses shared parameters across 128 layer calls, lifting a 1B backbone from 55.33 to 63.47 on MMStar and revealing a cross-loop Visual Aha Moment

The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
arXiv

Six-domain evaluation of GPT-6 Astra as an embodied policy: navigation leads, hybrid control lifts manipulation success, but direct in-hand control and dense-reference locomotion remain unreliable

This work systematically evaluates GPT-6 Astra as an embodied policy across six domains—gripper manipulation, dexterous manipulation, mobile manipulation, navigation, locomotion, and humanoid loco-manipulation—comparing direct control with hybrid control that cooperates with learned policies or whole-body controllers, finding leading navigation results (92% on RxR instruction following, 82% on HM3D object search), hybrid success of 48% on RoboDojo, 50% on DexJoCo, and 38.7% on RoboCasa365, while direct in-hand control and dense-reference locomotion remain unreliable, with substantial token use and inference latency recorded.
arXiv

Across 800 pretrained models, wild AI web text raises loss once data is plentiful, and a 31.1% AI share costs 1.6x the compute

Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
arXiv

ATLAS preserves relational geometry while calibrating the latent distribution in LeWM, improving goal-reaching success on PushT, TwoRoom, and OGBench-Cube, with the largest gain on higher-novelty TwoRoom episodes

The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
arXiv

Scaling synthetic pre-pretraining from 1B to 7B and 100B tokens keeps the token savings, but the authors trace them to long-range retrieval rather than a grammatical prior

The study systematically evaluates synthetic pre-pretraining (PPT) across four parameter scales from 500M to 7B, four pre-training data mixtures, five PPT tasks, and pre-training budgets up to 100B tokens, finding that downstream performance and token-efficiency gains persist at scale (saving at least 21B pre-training tokens at 3B) but that there is no consistent evidence the gains come from a grammatical prior; instead they arise from PPT tasks that improve long-range retrieval, are insensitive to code and math share, and diminish only when web text is absent.
arXiv

Removing the overlapping-window timing shortcut lets non-invasive brain-to-text reach 36.6% word error rate with five observations per word

The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
arXiv

SkillSeek matches an LLM-mediated retrieval loop on 89 SkillsBench tasks with BM25 plus a small reranker, cutting per-trial spend from USD 51.30 to USD 27.54

The work presents SkillSeek, an open-source two-stage skill retriever (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP), and across a grid of two skill pools and two backbones on the 89-task SkillsBench benchmark observes that plain bm25 alone records a pass rate at or above Liu et al.'s LLM-mediated loop on three of four settings, with a small cross-encoder covering the remaining difference on the fourth (34K pool with Qwen3.5) at the same 0.442, while total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the USD 27.41 no-skill baseline), a pattern the authors attribute to a first-stage recall ceiling.
Nature News

Danish foundation data show endometriosis research received EUR 173,958 versus EUR 254.9 million for diabetes, while a mouse study reports a tweaked niclosamide reversed lesion-induced macrophage changes

A Nature news story describes a mouse study, published in Advanced Healthcare Materials, in which researchers tweaked the antiparasitic drug niclosamide to target cells implicated in endometriosis, while an npj Women's Health analysis reports that among Denmark's top 100 grant-awarding foundations endometriosis research received EUR 173,958 compared with EUR 254,908,430 for diabetes and EUR 325,940 for inflammatory bowel disease, alongside a wide gap in media mentions.
arXiv

Endless Exam scores nine models on 14 families of math construction problems: GPT-6 Astra reaches 143.16 with tools, yet none of 30 published frontiers is surpassed

The authors introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems with 69 evaluation instances, where each submitted object is automatically checked for validity and given an uncapped relative quality score (100 marks average reference parity); across nine models and seventeen configurations, tool-free overall scores span 7.55 to 91.90, and with code and web access GPT-6 Astra, GPT-5.6 Luna and Claude Opus 5.5 reach 143.16, 119.35 and 180.06, while none of the 30 published-frontier references is surpassed.
arXiv

Galahad turns LLM document reading into a one-time cost with byte-exact KV memory: 100 of 100 on a 97,000-token corpus at 0.59–0.64 s per question

Sietse Schelpe of Corbenic AI presents Galahad, a memory layer in which Taliesin saves and reloads byte-exact KV state and Blaise passes the model only the section a question needs; on a recall test with 100 facts hidden in a 97,000-token corpus (Gemma 4 31B), Taliesin alone answered 98 of 100 on llama.cpp at 3.0 s and 572 J per question against 10 of 100, 9.3 s and 2,754 J without Galahad, and with Blaise added the model answered 100 of 100 on all three runtimes at 0.59–0.64 s and 200–213 J, while a tuned RAGFlow pipeline answered 77.
arXiv

Freezing agents and training only latent links: benign link training alone raises harmful compliance above text communication, and an RL attack lifts the mean from 27.9 to 76.9

This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.