AI Core
906 items
Across 800 pretrained models, wild AI web text raises loss once data is plentiful, and a 31.1% AI share costs 1.6x the compute
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Freezing a 7B backbone and training only 100K bias parameters: label-free TTRL reaches 76.67% on MATH-500 and transfers to 4,500 problems unseen during optimization
The work introduces label-free bias-only test-time reinforcement learning: the pretrained backbone stays frozen while only about 100K bias parameters (down_proj.bias across 28 MLP layers) are optimized against majority-vote pseudo-label rewards, reaching 76.67% on MATH-500 with Qwen2.5-7B and 79.50% with Qwen2.5-Math-7B, improving vision-language and audio reasoning under the same procedure, and transferring frozen bias vectors to 4,500 verified-disjoint MATH problems to lift Qwen2.5-7B from 46.3% to 70.9% and Qwen2.5-Math-7B from 52.5% to 75.4%, with pseudo-label reliability and accessible gradient energy explaining why a restricted subspace still adapts.
Teaching a Builder to design execution environments: meta-skills lift macro-average scores by 8.95 points with frozen model weights
The work introduces meta-skills — principles specifying when support is needed, what resources to provide, and how the Target should use them — which a Builder learns from the Target's execution feedback on a development set while both models' weights stay fixed, then freezes into a skill bank to construct harnesses for unseen tasks; across Harness-Bench and NewtonBench, full-bank meta-skills raise macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over delivering the same bank directly to the Target.
SUSTech and CityU propose Box2-Bench: frontier models use good workflows but lose up to 43.3 points under bad ones
The authors introduce Box2-Bench, which holds the model and task fixed while varying workflow availability and reliability across five matched conditions (none, good, bad, partial, mixed), finding that frontier models such as Gemini 3.7 Flash, DeepSeek V4 Flash, and GLM 5.2 benefit from good workflows in 8 of 12 model-task pairs but are hurt by bad workflows in all 12 (drops of 3.3 to 43.3 points), and that mixed workflows underperform their matched partial workflows in 10 of 12 pairs; training two open-weight models on bad workflows only, counterfactual SFT cuts the bad-workflow penalty from 20.0 to 6.7 points on AIME and from 9.8 to 0.6 points on WebShop while weakening use of held-out good workflows, and outcome-based RL restores AIME utilization from -3.3 to +6.
RIDE extrapolates RL teachers' hidden-state residuals and matches or beats the teacher on average across four base/teacher pairs
The work proposes RIDE (RL-Induced Direction Extrapolation): on student-generated trajectories it computes the layerwise hidden-state residual between an RL-trained teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets displaced beyond the teacher along that residual; across four base/RL-teacher pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) RIDE's mean Avg@16 on AIME24, AIME25 and AIMO approaches or exceeds the RL-trained teacher on every pair, the only compared method whose mean does so, and it consistently outperforms output-space extrapolation.
Soft Spatial Reasoning uses AdaptSoft to tune softness per step, lifting OmniSpatial weighted-average accuracy to 49.68
The work proposes Soft Spatial Reasoning, a post-training framework in which a large vision-language model forms a continuous soft state at each chain-of-thought step by mixing token embeddings instead of committing to a single token, with an AdaptSoft controller that adapts the degree of softness from the current hidden state and predictive uncertainty and is trained by a gradient-alignment objective; across OmniSpatial, SpatiaLab, and MindCube it reaches higher weighted-average accuracy than same-backbone hard-thinking and fixed-soft chain-of-thought baselines and the compared existing models.
RSIGame pairs local explore-diagnose-improve with a global best-checkpoint loop, lifting Qwen3.8-27B from 37.07 to 61.38 on Godot and past GPT-5.5's 50.26 one-shot score
RSIGame organizes automatic game development as recursive self-improvement: a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes issues, and performs evidence-grounded revision while an evolving checklist accumulates testing and improvement guidance; a global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression; and successful development experience is internalized into the generator through supervised fine-tuning, consistently improving game quality across 140 GameCraft-Bench tasks, two engines, and five generators under matched development budgets, with experience internalization enabling Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.
LoopVL reuses shared parameters across 128 layer calls, lifting a 1B backbone from 55.33 to 63.47 on MMStar and revealing a cross-loop Visual Aha Moment
The authors introduce LoopVL, extending looped Transformers to vision-language models: the language backbone is pretrained from scratch on HRM-Text's open-source framework and data, using nested Module-Loop and Model-Loop computation over shared L/H parameter stacks, with the default H2L3 configuration executing 128 Transformer-layer calls per forward pass; under the same 0.14T-token training budget, LoopVL outperforms a same-depth Transformer-VL 1B baseline on MMStar, RealWorldQA, and ChartQA and approaches or surpasses larger 4B dense baselines, while the authors report cross-cycle visual attention reallocation they call a Visual Aha Moment.
EviRover teaches a 4B model to look beyond a glance: 30-point average gain on EviLens and 15 points on BrowseComp-VL
The work defines perception under insufficient evidence, reformulates perception tasks such as grounding, segmentation and counting as an agentic evidence-seeking process, builds two data generation pipelines yielding EviRover-SFT-5K and EviRover-RL-12K plus the human-verified 688-instance EviLens benchmark spanning five perception categories, and shows that a 4B model trained with supervised fine-tuning followed by agentic reinforcement learning improves over its Qwen3-VL-4B-Instruct backbone by 30 points on average on EviLens, reaches performance comparable to advanced proprietary models, and transfers to WebEyes, ReasonSeg, RefCOCOg, MMMU, MMMU-Pro, MathVerse and BrowseComp-VL with a 15-point gain.
Page 5 · showing 10