AI Core
906 items
Tracking fine-tuning trajectories shows safety-aligned LLMs are not steered toward harm but lose their safe subspace, and projecting out the harmful gradient subspace cuts free-generation emergent misalignment by up to 80.0% on Qwen2.5-14B-IT
The work presents a dynamic second-order geometric study of emergent misalignment (EM): tracking training trajectories shows directional Hessian curvature concentrates on semantic pivot tokens, Grassmannian projections show the harmful-safe gap widens mainly because safe-gradient overlap declines rather than through rotation toward a new malicious circuit, and on this basis it introduces a parameter-level Geometric Mitigation Framework that orthogonally projects the empirical harmful gradient subspace out of LoRA updates, suppressing free-generation EM by up to 80.0% on Qwen2.5-14B-IT and, across four open-weight instruction-based model families (3B–20B), teacher-forced evaluation shows the same harmful subspace controls the conditional support of frozen EM responses.
DuoOPD Rewrites Distillation Feedback from Joint Teacher-Student Outcomes, Lifting Macro Accuracy by 2.58 Points on Qwen3 and 5.98 on Llama over OPD
The work introduces DuoOPD, a multi-task on-policy distillation method in which the student's verified outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it—when only the teacher succeeds its verified answer becomes context for scoring the student's failed response, and when only the student succeeds a weight shared within the task reinforces the whole response, with one rule covering all four outcome combinations and no task-specific settings; across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation.
Labeling agents with their model family split multi-agent groups along the label, raising rounds and tokens and cutting success from 96% to 81%
Across two no-stakes cooperative games (Leader Election and Exclusion) and the GPQA-Diamond reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families, the study finds that when agents know each other's model family the interaction graph partitions into factions along that label, that the split follows the visible label rather than the underlying architecture (it persists under shuffled labels and neutral color tags and disappears when labels are removed), and that in strictly cooperative tasks labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision while success drops from 96% to 81%.
CoEvoWhen lets a frozen VLM coevolve policies and tools for ultra-long video, raising ExtremeWhenBench mIoU by 74.9% while cutting visual tokens by 11.4%
CoEvoWhen introduces a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories into a reusable skill without updating model parameters, improving ultra-long video temporal grounding accuracy and reducing inference visual token cost across five benchmarks and three VLMs, with the evolved skill transferring to general long-video QA without additional task-specific evolution.
AdviSD trains a Qwen3-8B advisor to learn selectively from feedback corrections, beating advisor-GRPO by 4.2–6.4 points on BFCL-v3
The work introduces AdviSD, in which a small trainable advisor (Qwen3-8B) steers frozen frontier executors (Gemini 3.7 Flash, Claude Sonnet 4.6) with natural-language advice by pairing outcome-based GRPO with selective feedback-conditioned self-distillation: reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of that difference to choose which decisions to supervise, reaching 4.2–6.4 percentage points above advisor-GRPO on BFCL-v3 and 3.9–5.1 score points above it on EnvScaler while beating matched-count random selection and a no-gate variant.
Freezing agents and training only latent links: benign link training alone raises harmful compliance above text communication, and an RL attack lifts the mean from 27.9 to 76.9
This work studies the safety of latent communication in multi-agent systems, where lightweight trainable links replace text messages, and finds that with all agent parameters frozen, benign link training alone increases harmful compliance relative to text communication, that an attacker can amplify this via supervised optimization on harmful query-response pairs, poisoning benign training data, or a reward-guided reinforcement-learning attack that needs no harmful target responses, raising the mean harmful-compliance score from 27.9 with benignly trained links to 76.9 across three topologies and four safety benchmarks, and that adapting the rewards toward safer behavior repairs compromised links by updating only the links.
Physis-Lang rewrites video world models with self-evolving physical language: open-source Cosmos3-Nano surpasses Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench
Physis-Lang treats physical language as a shared, optimizable representation across data curation, model training and video generation, using a physics-aware critic and PhysCapBench to drive an agentic loop that refines physical captioning instructions and language-guided retrieval to cover missing physical processes, yielding consistent gains on four physical video benchmarks with Wan and Cosmos backbones, where models built on the open-source Cosmos3-Nano surpass Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench.
WorldAuditBench tests multimodal agents on 213 3D world-auditing tasks: best model 42.3% vs. humans 83.4%
The authors built WorldAuditBench, a benchmark of 213 tasks across 13 Unreal Engine 5 and Three.js environments and five anomaly families for interactive 3D world auditing, and compared a single-model VLM agent with a two-stage VLA–VLM auditor, finding VLM agents reach 28.2%–42.3% success versus 6.6%–17.4% for two-stage auditors, both far below the 83.4% human rate.
DAGent grows its task DAG batch by batch from node-level evidence, beating the strongest open-source baseline by 5.3/5.8/2.0 points on BrowseComp-Plus, GAIA and xbench-DeepSearch
DAGent introduces Evaluate-then-Grow incremental planning, in which an Orchestrator grows the task DAG one batch at a time conditioned on confidence and uncertainty signals from completed nodes, paired with a hierarchical context layer that propagates compact QueryDocs by default while preserving full execution traces for on-demand recall, and with DAGRPO, a DAG-conditioned RL adaptation combining topology-conditioned credit with structural compliance regularization; it surpasses the strongest open-source baseline by 5.3/5.8/2.0 points on BrowseComp-Plus, GAIA and xbench-DeepSearch at the Qwen3-235B-A22B scale, improves over a same-budget outcome-only GRPO baseline by 3.
Page 8 · showing 10