Skip to main content

Search

“All disciplines” · 1367 results

Page 7 · showing 20
arXiv

Tracking fine-tuning trajectories shows safety-aligned LLMs are not steered toward harm but lose their safe subspace, and projecting out the harmful gradient subspace cuts free-generation emergent misalignment by up to 80.0% on Qwen2.5-14B-IT

The work presents a dynamic second-order geometric study of emergent misalignment (EM): tracking training trajectories shows directional Hessian curvature concentrates on semantic pivot tokens, Grassmannian projections show the harmful-safe gap widens mainly because safe-gradient overlap declines rather than through rotation toward a new malicious circuit, and on this basis it introduces a parameter-level Geometric Mitigation Framework that orthogonally projects the empirical harmful gradient subspace out of LoRA updates, suppressing free-generation EM by up to 80.0% on Qwen2.5-14B-IT and, across four open-weight instruction-based model families (3B–20B), teacher-forced evaluation shows the same harmful subspace controls the conditional support of frozen EM responses.
arXiv

EvoDuet lets LLMs retrieve the web by knowledge gap during evolutionary search, raising OpenEvolve's normalized discovery gain from 74.1% to 78.0% (GPT-5.6-Luna) and from 61.3% to 82.3% (Gemini-3.8-Flash) across 21 optimization tasks and surpassing prior best scores on eight

The work introduces EvoDuet, a bi-level method that co-evolves solutions and web search queries with fixed model parameters: at each iteration a knowledge-gap-based retrieval gate decides whether to retrieve new documents, reuse stored ones, or proceed without them, an inner loop refines queries and ranks documents by the solution scores they are predicted to yield, and an outer loop generates candidates in parallel and records evaluated outcomes for later searches; across 21 optimization tasks with one candidate per iteration it raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, Qwen3.
arXiv

ThinkV2V makes MLLMs think before editing: a 5B model tops both complex and standard video-editing benchmarks

The work proposes ThinkV2V, a reasoning-driven video-editing framework that explicitly activates multimodal large language model thinking before visual generation: it uses Qwen3-VL-Thinking-8B to perform chain-of-thought reasoning over the source video and instruction and output a refined prompt, injects the hidden states into a Wan-2.1-5B DiT through a learnable-query connector, pairs this with Progressive Curriculum Training and Inference-Time Thinking Scaling, and builds the ThinkV2V-150K dataset and ThinkV2V-Bench; under two judge models, the 5B-scale model achieves the best overall scores on both complex and standard editing scenarios and surpasses several 10B-scale baselines.
arXiv

Selecting distillation positions by entropy shift lifts LLM judges 2–9 points over Dr. GRPO on subjective tasks

This work studies training LLM judges from natural language feedback, proposes using the per-position entropy shift between teacher and student to distinguish context sharpening from context spreading, and masks high-entropy-shift positions so that self-distilled judges outperform outcome-supervised RL (Dr. GRPO) by 2–9 percentage points on the evaluated subjective subcategories while remaining competitive on objective ones.
arXiv

OASIS restricts self-distillation supervision to verified trajectories, widening its average gain over OPSD from 0.59 to 3.05 points across Qwen3-1.7B, 4B and 8B

Using truncated-continuation probes, the work finds that in on-policy self-distillation (OPSD) the teacher's advantage over the student is concentrated on failed trajectories (77% versus 37% recovery at 1.7B) and shrinks with scale, and therefore proposes OASIS: keeping the OPSD objective unchanged, it applies supervision only along the shortest outcome-verified on-policy trajectory and replaces the written reference solution with a distinct same-problem self-generated rollout, typically the student's own unverified attempt, improving mean Avg@12 over OPSD by 0.59, 1.86 and 3.05 points on Qwen3-1.7B, 4B and 8B across AIME 2024, AIME 2025 and HMMT 2025 while requiring no written solutions.
arXiv

MIST stress test: adding an image shifts about a fifth of VLM judge labels, regardless of what the image shows

The authors built MIST (the Misleading-Image Stress Test), 200 English sentences each containing a phrase readable either figuratively or literally and shown with an aligned image, a misleading image, or no image; across thirteen VLM judges an aligned image changed 20.5% of labels and a misleading one 19.4%, versus 11.6% when only the ignore-the-image instruction was deleted with the image left in place, and only 37% of the labels that differ between the two images moved toward the sense shown, while agreement with the human annotators was unchanged whether the image was absent, aligned, or misleading.
arXiv

NavHarness carries maps, search records and house notes into the next conversation, lifting GOAT-Bench s-SR by 18.6 to 34.8 points

NavHarness is a training-free embodied harness that makes memory processing part of the navigation loop: each task opens a fresh multi-round agentic session that inherits prior experience through maps, task records and house notes, while outcome verification and run-end consolidation decide what later sessions inherit; on GOAT-Bench it improves s-SR over context-only independent sessions by 18.6 points with GPT-6 Astra, 22.6 with Opus 5, 34.8 with GPT-4o and 30.3 with Qwen3.8-27B, and with SLAM-estimated poses Astra reaches 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE.
arXiv

AREX-2 trains a 27B agent on long-horizon reflective trajectories, reaching 81.8 on MLE-bench Lite and 84.0 on BrowseComp

AREX-2 defines agent self-improvement as turning more test-time rounds into a better solution, synthesizes long-horizon improvement trajectories from machine learning engineering and algorithmic programming tasks that retain failures and regressions, and after fine-tuning Qwen3.8-27B reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS while lifting BrowseComp to 84.0, HLE to 52.6, GAIA to 92.2 and DeepSearchQA to 93.8 without adding new deep-research data, with continued gains as the round budget grows.
arXiv

Keeping only the top-attention tokens without retraining, nine checkpoints need effective attention sets of 7 to 283 tokens, and longer context raises the required set size

Without retraining, this work retains only the tokens with the highest attention weights at each head, layer, and query while keeping their original weights unchanged, and estimates the effective attention set size needed to stay within a chosen loss tolerance by measuring the increase in negative log-likelihood (NLL); it finds that relatively small selected sets keep NLL close to the full-attention baseline, that attention-based selection substantially outperforms random selection, that the required set size grows with context while its fraction of context decreases, that in BABILong experiments with a fixed annotated supporting fact additional background pushes support tokens down the attention ranking and reduces their attention mass, and that renormalizing the retained weights can subs
arXiv

Splitting masked-diffusion decoding into five axes, researchers find state adaptation pays off only in a few predictable states, and selective intervention lifts task-macro utility by 3.2, 4.0, and 5.6 points across three models

The work factorizes masked diffusion language model (MDM) inference into five axes—score, cardinality, region, commitment, and planning—defines adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action, finds across three models (LLaDA-8B-Instruct, LLaDA-1.5, Dream-7B) and ten tasks that adaptation opportunities are highly heterogeneous, and proposes selective adaptation in which lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing for example 56.9% of the candidate-set oracle opportunity by adapting only the top 10% of states on LLaDA-8B constrained JSON filling, with end-to-end selective decoding raising task-macro utility from 0.392 to 0.424 (LLaDA-8B), 0.344 to 0.
arXiv

Physis-Lang rewrites video world models with self-evolving physical language: open-source Cosmos3-Nano surpasses Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench

Physis-Lang treats physical language as a shared, optimizable representation across data curation, model training and video generation, using a physics-aware critic and PhysCapBench to drive an agentic loop that refines physical captioning instructions and language-guided retrieval to cover missing physical processes, yielding consistent gains on four physical video benchmarks with Wan and Cosmos backbones, where models built on the open-source Cosmos3-Nano surpass Veo 3.1 on Physics-IQ Verified, PhyGround and PhyGenBench.
Nature News

Two years after the Anthropocene proposal was rejected, scholars allege the vote was not transparent and call for the documents to be released, while a review of 12 global stratigraphic records reaffirms 1952 plutonium as a viable boundary marker

In 2024 the International Union of Geological Sciences upheld the rejection of a proposal to establish the Anthropocene as a formal geological epoch (a stratigrapher subcommittee voted 12 against, 4 in favor, 2 abstaining); two years later, Jürgen Renn and co-authors argue in Earth's Future that unpublished documents show the rejection was illegitimate and its rationale never adequately explained, calling on the IUGS to release key documents and submit the process to independent review, while supporters publish a review in Nature Reviews Earth & Environment compiling multi-proxy evidence from 12 globally distributed stratigraphic records showing mid-twentieth-century Earth system changes are abrupt, globally synchronous and stratigraphically distinct, with a sharp 1952 plutonium increase f
arXiv

DuoOPD Rewrites Distillation Feedback from Joint Teacher-Student Outcomes, Lifting Macro Accuracy by 2.58 Points on Qwen3 and 5.98 on Llama over OPD

The work introduces DuoOPD, a multi-task on-policy distillation method in which the student's verified outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it—when only the teacher succeeds its verified answer becomes context for scoring the student's failed response, and when only the student succeeds a weight shared within the task reinforces the whole response, with one rule covering all four outcome combinations and no task-specific settings; across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation.
arXiv

Cross-meeting speaker attribution: ThyVoice leads all evaluated commercial cascades at a 47.13 mean SI-cpWER, while per-meeting leader ElevenLabs gains about 30 percentage points of attribution error on CHiME-6

The work introduces SI-cpWER, a metric that scores cpWER under one corpus-global speaker-ID map, and evaluates five commercial diarize-then-identify cascades, two open baselines, and the end-to-end reference system ThyVoice on CHiME-8 NOTSOFAR (clean and noise-augmented) plus CHiME-6, finding that requiring one persistent identity across meetings changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions, with a full-panel mean of 47.13 versus 54.75 for the next system.
arXiv

S²D-OPD keeps only the top 10% of states per response by teacher–reference JSD, lifting held-out accuracy from 46.36% to 47.31% in seven of eight teacher–student settings

The work shows that Direct-OPD's token-level log-ratio reward measures only relative change and can stay fixed while the probability mass that both checkpoints assign to the student's candidate tokens vanishes, whereas the teacher–reference JSD and both KL directions vanish with that mass; it therefore proposes S²D-OPD, which ranks student-sampled states by teacher–reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% per response, improving held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings across two teacher pairs and four student models from 1.7B to 8B parameters and matching it in the eighth, without extra forward passes.
arXiv

PatchHolmes lets an agent read all 100 candidate commits at once, lifting patch-retrieval Recall@1 from 32.63% to 59.95%

PatchHolmes is a two-phase patch retrieval system whose first stage fuses BM25 with time decay and Qwen3-Embedding-8B dense retrieval via RRF into a top-100 candidate set, and whose second stage runs a frozen open-weight LLM agent that views all 100 candidates listwise through four tools (list_candidates, read_commit, read_file_diff, submit_answer) and submits a single best commit; on 809 CVEs from GitHubAD it reaches 59.95% Recall@1, 25.34% above the pointwise classifier Favia and 31.40% above IRCoT, adds 27.32% Recall@1 over the retriever's top pick on identical candidates, and transfers unchanged to PatchFinder_top10 to lift Recall@1 from 24.28% to 39.86%.
arXiv

ReaLVR supervises latent visual reasoning with visual evidence, reaching a 63.7% five-task average on Qwen2.5-VL-7B and scaling latent visual reasoning to 235B

The work first analyzes latent-token behavior in latent visual reasoning (LVR) and identifies a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that alter the correct answer, which the authors attribute to the lack of explicit supervision during GRPO training; it then proposes ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories, using correct-versus-wrong answer contrast to decide where stronger supervision is needed and relevant-versus-mismatched visual evidence to specify what to preserve, consistently outperforming evaluated LVR baselines across three model families, reaching a 63.7% five-task average on Qwen2.5-VL-7B, and being the first to scale visual reasoning in latent space up to 235B parameters.
arXiv

Imagine3D-LLM has multimodal LLMs imagine a compact 3D scene before answering, reaching 68.5 average on SPAR-Bench at 7B scale

The work introduces Imagine3D-LLM: a small set of learnable Gaussian summary tokens is appended after the image tokens, decoded into a compact 3D Gaussian Splatting representation, and trained jointly with a photometric reconstruction loss and standard next-token prediction, so that a multi-view multimodal LLM assembles a coarse 3D scene representation before answering; it consistently outperforms prior approaches across seven spatial reasoning and 3D scene understanding benchmarks, reaching 68.5 average on SPAR-Bench and surpassing the strongest spatially-aware baseline 3DThinker-7B by 5.2 points.
arXiv

RL post-training of vision-language-action models concentrates parameter updates in Timestep Modules that hold only 14.83%–27.58% of the action expert, and their low-rank directions predict and improve task success

This work systematically analyzes how reinforcement learning (RL) reshapes flow-based vision-language-action (VLA) models, finding that across π0.5 and GR00T N1.5/N1.6 on LIBERO, ManiSkill, MetaWorld, and CALVIN, RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules, which hold 14.83%–27.58% of the parameters yet capture a disproportionate share of the RL gain, with shift-vector update directions predicting task success at up to 99.6% ROC-AUC and steering along them improving policies without additional RL training.
arXiv

EVOKE ranks the same actions under alternative goals at fixed states, pushing pretrained world knowledge into decisions: ALFWorld unseen success rises from 60.4% to 91.8%

EVOKE is a post-training method that holds the environment state and interaction history fixed while introducing alternative goals for the same candidate actions and training on contrastive rankings, forcing the policy to elicit action-consequence knowledge already present from pretraining; across ALFWorld, WebShop, and search-based QA on three backbones it achieves the best average, raising ALFWorld unseen-game success from 60.4% for πboot to 91.8% while using fewer actions.