Skip to main content

Search

“All disciplines” · 1409 results

Page 15 · showing 20
arXiv

Q&D Trains an 8B Questioner on the Consequences of Its Questions: Required-Evidence Coverage on MuSiQue Rises from 78% to 90%, and Retail Task Success from 13% to 34%

The work adds a content axis to agent proactivity—horizontal proactivity pursues unstated information the current context already identifies, vertical proactivity pursues needs only earlier evidence reveals—scores runs against a need graph recovered mechanically from multi-hop benchmarks' own decompositions with no model judge, and proposes Q&D (questioner and drafter), which forks a recorded run at one step, samples eight candidate questions, continues each to the end, and prefers the question whose continuation retrieves more required evidence, so no reward model or judge is needed; on held-out splits of MuSiQue, StrategyQA and 2WikiMultiHopQA, at equal retrieval spend the trained 8B questioner recovers more required evidence than the same model prompted (89.5% versus 78.
arXiv

EVO-WAM lets world action models self-train on their own generated video, lifting RoboTwin unseen-task success from 26.9% to 68.0%

The work proposes EVO-WAM, a framework that enables complete autoregressive rollouts without external execution feedback via state prediction and anchored multi-frame context, then selects task-completing prefixes with a vision-language model and verifies video-action consistency with an inverse dynamics model for iterative self-training; on seven unseen RoboTwin 2.0 tasks it raises average success from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, and on three real-world long-horizon composite tasks it raises Cosmos3 from 20.0% to 76.7%.
arXiv

AnisoWM swaps isotropic regularization for a learnable diagonal covariance target and beats LeWM's planning success in all four visual control environments

The work shows that isotropic Gaussian regularization (SIGReg) can make Euclidean latent planning cost rank feasible outcomes differently from the task cost, and introduces AnisoWM with ΛReg, which learns a diagonal covariance target under fixed-trace and anisotropy constraints while leaving the prediction objective and Euclidean planner unchanged, improving planning success over LeWM in all four visual control environments (TwoRoom, Reacher, PushT, Cube) and yielding better agreement between latent costs and task outcomes.
arXiv

ReImaGin uses image generation as a multimodal reasoning tool, beating text-only reasoning and specialist vision-tool baselines by up to 25% across six visual reasoning tasks

The work proposes ReImaGin, a training-free multimodal agent framework in which a multimodal LLM calls a free-form generate_image tool backed by an instruction-following image generation model inside its reasoning loop; across six tasks—depth perception, puzzle completion, occlusion counting, collision prediction, multi-view spatial reasoning, and path tracing—it consistently outperforms text-only reasoning and the fixed specialist vision-tool baseline Visual Sketchpad, with gains of up to 25% on path tracing and about 40% relative on occlusion counting, and it further shows that effective visual reasoning strategies can be discovered automatically via prompt optimization.
arXiv

Tsinghua's Leap Lab finds in a matched comparison that world action models generalize from a single inference-time forward pass, not from denoising a clean future

Under matched backbone, data, and training budget, the study compares explicit and latent world action models and finds that latent models match explicit ones in distribution (96.85 vs 97.75) yet fall behind on all three generalization axes—environmental perturbation, data efficiency, and task generalization—while leaving future video tokens at pure noise and running a single forward pass recovers most of the gap, motivating Simple-WAM, which leads explicit models in generalization at latent-level inference speed.
arXiv

One token-level interaction decomposition links data selection, collision-versus-erosion forgetting, and plasticity loss into a single learning-dynamics account

The work derives a token- and layer-wise decomposition of how learning from one token changes another prediction, splitting it into a direct force-alignment channel (CH1) and a diffuse channel mediated by shared readout geometry (CH2), with a forward-computable approximation using only logits, hidden states, and the readout; following this interaction over time, the authors report that positive interaction supports data selection, negative interaction splits into concentrated collision and accumulated erosion with distinct controls, and that declining task-conditioned readout transmission during long-horizon training tracks declining future learnability, with readout restoration partially recovering plasticity.
arXiv

Sparse crosscoders show on-policy distillation adds no new student features but reweights features the student already shares with the teacher

Using sparse crosscoders and a proposed swap readout, the study reads how each student checkpoint uses every feature before and after on-policy distillation (OPD) across three OPD settings, finding that OPD neither creates the student's own features nor passes on the teacher's, but slightly reweights features the student already shares with the teacher (over 98% of frequently used features change firing rate by less than 20%), with the largest changes concentrated on decision tokens such as Wait, Hmm, and So; the SFT warm-up on the teacher's rollouts that commonly precedes OPD also adds no features, reweighting shared features partly along OPD's direction and partly beyond it, and imposing this reweighting on a directly distilled student's features without changing its weights brings its a
arXiv

Fudan team trains Chinese-Jev on 10 million Chinese decisions, reaching 69.20% general accuracy and roughly 20x faster than Jev

The work introduces Chinese-Jev: a unified data pipeline converts heterogeneous Chinese annotations into probability targets over candidate options, a lightweight encoder-only model is first pre-trained on 10 million general Chinese decisions and then fine-tuned separately for medicine, law, and finance, and CJ-Bench is released with 307,900 held-out decisions; after pre-training it reaches 69.20% accuracy on general tasks, a 1.24% relative gain over the closed-source Jev, cuts expected calibration error from 11.45% to 3.78% at 14 ms latency, improves medical accuracy over Jev by 4.0%, reaches 92% of Jev's average accuracy across the three specialized domains at 15 ms average latency, and an INT8-quantized model runs about 1 second per decision in a mobile browser.
arXiv

TGRL turns temperature differences into a training signal: +1.6% math average, +196.7 CodeForces rating, with no extra rollout budget

The work proposes Temperature-Grouped Reinforcement Learning (TGRL), which partitions each prompt's rollout group into low- and high-temperature subsets, estimates prompt-level exploration gain from their reward contrast, and allocates that gain as token-level credit using Jensen–Shannon divergence between the temperature-scaled next-token distributions induced by the same logits; across 11 benchmarks, TGRL improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%, reaching equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget.
arXiv

KAIST team turns prompt-template disagreement into preference supervision, lifting open-vocabulary segmentation across the MESS benchmark without pixel-level labels

The work proposes a preference-guided adaptation framework for open-vocabulary semantic segmentation that converts the segmentation differences different prompt templates produce on the same image (prompt disagreement) into binary preference supervision, adapting models with Region-Localized Preference Optimization (RLPO) plus consistency regularization in a streaming single-step setting, achieving consistent gains on the five domain groups of the MESS benchmark across SAN and CAT-Seg backbones (ViT-B/16 and ViT-L/14) without any pixel-level annotation and remaining effective under noisy preferences.
arXiv

GRAFT swaps trajectories between two heterogeneous models, beating GRPO for both at equal budget with a 2.1-point average gain

The work proposes GRAFT, a framework that in Reinforcement Learning with Verifiable Rewards (RLVR) replaces a receiver's all-fail rollout groups with peer groups containing both successful and unsuccessful responses, controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping; across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT improves both models over GRPO at the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points, with 1.8 points on average preserved when reusing stored peer trajectories.
arXiv

NVIDIA turns distillation into pluggable LoRAs: train once, then deploy to 54 video models without retraining

LongLive-Plug introduces a once-for-all distillation framework that distills single-pass classifier-free guidance, four-step sampling, and long-context error correction into reusable LoRAs on a base model, enabling training-free plug-and-play deployment to 54 downstream video models across three backbone families, with results on SCOPE and Wan2.2-Fun-5B-Control that beat naive four-step sampling and approach per-target distillation.
arXiv

Porcedda's Sys1Cal-v1 shows Jev's Choice probabilities are systematically distorted, and restoring an uncertainty component lifts median soft accuracy from 0.771 to 0.978

Porcedda built Sys1Cal-v1, a synthetic dataset of 365 True/False items derived from 92 probability problems whose proposition probabilities are known by construction, and used total variation distance and distributional overlap (soft accuracy) to evaluate Jev's Noul, Choice, and Score primitives plus the open-source baseline SemIf, finding that Jev's Choice probabilities are systematically distorted and that inverting a latent third truth value of uncertainty estimated from Score expectations raises median Choice soft accuracy from 0.771 to 0.978.
arXiv

OmniTaskonomy maps 19 generation tasks against 25 understanding capabilities: I2I-then-I2T training improves understanding, and gradient alignment tracks the gains

Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.
arXiv

Shared or dual projections? A bias–variance boundary gives the criterion, and CARS cuts held-out regret by 49–96% across five datasets

The work develops a bias–variance theory for low-rank bilinear scoring in dense retrieval, proving that dual projections have lower risk exactly when squared directional signal exceeds the estimation cost of their extra degrees of freedom, and uses it to build the cross-fitted CARS selector, which reaches 90.1% mean geometry-selection accuracy and reduces held-out regret by 49–96% relative to the better fixed geometry across multiple datasets and embedding models.
arXiv

Raven composes model–harness pairs into a multi-agent ecosystem and leads planning, four specialist domains, and skill reuse over Claude Code and other baselines

Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.
arXiv

ByteDance team finds chunked KV-cache compression makes long-context retrieval periodically weak at the compression stride, with up to 40 percentage points between phases in DeepSeek-V4

The work names and systematically characterizes phase sensitivity in models with chunked KV-cache compression: the same information is easier or harder to retrieve depending on its phase relative to compression-window boundaries, with long-context retrieval accuracy in the DeepSeek-V4 family differing by up to 40.2 percentage points across phases at 128K context; the authors reproduce the effect by pretraining 23 compressed models from scratch (Qwen3-0.6B backbone, 100B tokens), use KV-head and layer mean-replacement knockouts to reveal phase specialization, and prove in an idealized induction model that gradient flow drives compression gates toward static preferences, showing that average benchmark scores can conceal periodic positional failures.
arXiv

Ant Group proposes Marathoner: synthesizing tasks from million-line-scale GitHub PRs lets a 9B open model work 10+ hours and make 1000+ tool calls

The work proposes Marathoner, an autonomous agentic model trained through a post-training pipeline of ultra-long-horizon task synthesis, rejection sampling finetuning, and reinforcement learning, achieving consistent and substantial improvements over the base Qwen3.5-9B across 5 ultra-long-horizon benchmarks and surpassing strong proprietary models on some, with analysis showing it can execute for 10+ hours and make 1000+ tool calls on highly challenging tasks.
arXiv

Ranking positions by chat structure lets auditors explain 5% of tokens and keep nearly all threat-detection success

Across about 4.7 million natural language autoencoder (NLA) explanations on four models and four auditing datasets, the study compares thirteen position-scoring signals computable in a single forward pass against a ranker trained only on chat structure, finding that chat structure beats the best individual signal in twelve of fourteen model-dataset cells, that explaining just 5% of positions retains nearly all of the success rate from explaining every position on three of four datasets, and that pretrained verbalizers recover words a fine-tuned model learned to conceal without additional verbalizer training.
arXiv

AutoDataBench makes agents deliver training tasks one at a time: five frontier agents all score below 20 out of 100 at 45 minutes, with difficulty calibration rather than mode coverage as the binding constraint

The work introduces AutoDataBench, which turns "can an agent write a training task that a data pipeline would accept sample by sample" into an evaluation target: given an original benchmark task and a record of the target model attempting it, an author agent must write a new task for the same suite, and the score multiplies a gate, whether the target model's pass rate lands in a band, and coverage of a hidden rubric of behavioural modes; across three benchmarks of executable agent tasks, no agent evaluated scores above 20 out of 100 at the default 45-minute budget, and giving the strongest agent four times as long raises the score substantially while the cost of one usable task stays almost unchanged.