Published AI and science developments
Across 800 pretrained models, wild AI web text raises loss once data is plentiful, and a 31.1% AI share costs 1.6x the compute
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
More selected stories
Freezing agents and training only latent links: benign link training alone raises harmful compliance above text communication, and an RL attack lifts the mean from 27.9 to 76.9
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
Turning context into an editable file: CLM lets models manage their own context, raising BrowseComp-Plus accuracy by 11.4% with 21.5% fewer FLOPs
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.
YuE2 writes a readable score before rendering, and experts prefer symbolic planning 49.3% to 34.6% on overall quality
YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
ZJU team releases EMem-Bench: across 2,554 long-horizon embodied episodes the strongest model Gemini-3-Flash reaches only 64.2% average success, while its EMem memory system lifts GPT-5.4-mini by 23.6 points
The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
Nereus adapts RL post-training parallelism on the fly: 27.7% lower average step latency on a real-data trace and up to 7.27x OpenRLHF throughput for 8B PPO
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Physis-Lang rewrites video world models with self-evolving physical language: open-source Cosmos3-Nano surpasses Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench
Physis-Lang treats physical language as a shared, optimizable representation across data curation, model training and video generation, using a physics-aware critic and PhysCapBench to drive an agentic loop that refines physical captioning instructions and language-guided retrieval to cover missing physical processes, yielding consistent gains on four physical video benchmarks with Wan and Cosmos backbones, where models built on the open-source Cosmos3-Nano surpass Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench.
Tracking fine-tuning trajectories shows safety-aligned LLMs are not steered toward harm but lose their safe subspace, and projecting out the harmful gradient subspace cuts free-generation emergent misalignment by up to 80.0% on Qwen2.5-14B-IT
The work presents a dynamic second-order geometric study of emergent misalignment (EM): tracking training trajectories shows directional Hessian curvature concentrates on semantic pivot tokens, Grassmannian projections show the harmful-safe gap widens mainly because safe-gradient overlap declines rather than through rotation toward a new malicious circuit, and on this basis it introduces a parameter-level Geometric Mitigation Framework that orthogonally projects the empirical harmful gradient subspace out of LoRA updates, suppressing free-generation EM by up to 80.0% on Qwen2.5-14B-IT and, across four open-weight instruction-based model families (3B–20B), teacher-forced evaluation shows the same harmful subspace controls the conditional support of frozen EM responses.
VoxParity tests 28 voice agents on 183 scenarios: when audio calls for protection, every system more often executes the routine request (41% against 12%)
The work introduces VoxParity, a benchmark of 183 scenarios from 14 sectors in which one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper) and the correct executable tool call changes with it; across 206 cue-bearing cells, all 28 systems (including nine production realtime agents) execute the routine request on protective calls at 41% against 12% over-triggering on clean calls, and only 11 of the 23 systems with a transcript path pass the words-only null test, with passes coming almost entirely from items that state the rule.
VoxMem tests 15 audio LLMs on 3,196 questions: none tops 40% at 32K, and swapping audio for transcripts drops speaker accuracy from 69.8% to 10.3%
The authors propose a two-axis taxonomy of spoken conversational memory — acoustic evidence type (speech semantics, speaker identity, paralinguistic cues, environmental sound) crossed with memory operation (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal) — and build VoxMem on it: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) at four context budgets from 8K to 64K tokens; evaluating 15 large audio language models, no model exceeds 40% overall accuracy at 32K (best 38.5%), models remember what was said far better than who said it, how it was said, or what was audible, and replacing audio with exact transcripts drops speaker accuracy from 69.8% to 10.3% while speech-semantics accuracy barely moves (75.9% to 71.0%).
AgentTell benchmark shows browser-use agents leak private user information through click choices in 61.1% of sessions, and falsely assure users of privacy in 34.5% of leaking sessions
The authors define and formalize behavioural side-channel leakage in browser-use agents, introduce the AgentTell benchmark of 20 scenarios and 100 tasks, and evaluate six backbones across 9,760 sessions, finding that agents carrying a secret reveal it through task-directed action choices in 61.1% of sessions despite an explicit privacy instruction, that agents still leak in 56.7% of sessions where their own memory states the secret must not be shared, and that in 34.5% of leaking sessions their final responses falsely assure users that no information was disclosed.
Across 800 pretrained models, wild AI web text raises loss once data is plentiful, and a 31.1% AI share costs 1.6x the compute
Using EditLens and Pangram to label Common Crawl web data from 2021 to 2026, the authors find that 27.5% of tokens passing FineWeb quality filtering in June 2026 are AI-generated, rising to 31.1% in August, and by pretraining 800 models from 19.9M to 973M parameters while varying the ratio of added AI to human tokens they fit a new scaling law with separate benefit and harm terms: AI tokens lower loss for data-starved models before saturating and reversing into harm, raise loss almost immediately for Chinchilla-optimal models, and the law predicts held-out sizes with lower error than eleven existing laws, implying that training at August 2026's AI share takes 1.6x the compute.
Freezing agents and training only latent links: benign link training alone raises harmful compliance above text communication, and an RL attack lifts the mean from 27.9 to 76.9
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
EngiWorld tests 1,301 real engineering tasks and finds the strongest model reaches only 44.3 EngiScore, with 3.6% success on multi-software attempts
The authors built EngiWorld, a benchmark structured around the complete engineering design loop, with 1,301 expert-curated tasks across 6 domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, evaluated through a unified domain-verifier suite that checks geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts; across seven frontier models the strongest reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed.
Turning context into an editable file: CLM lets models manage their own context, raising BrowseComp-Plus accuracy by 11.4% with 21.5% fewer FLOPs
The work introduces Context Language Models (CLMs), which mirror a model's live context as a freely editable file that the model can rewrite with Bash, shifting context management from external harness control to an intrinsic model capability; applied zero-shot to existing models such as Qwen3.6-27B and GPT5.6-Sol, CLMs achieve 11.4% higher accuracy with 21.5% fewer prefix-reuse FLOPs than the strongest baseline on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater end-to-end speedup at the same compute on a 24-hour multi-repository agent-swarm task, and can further learn context-management strategies through natural-language instructions, textual evolution, and online reinforcement learning.