Skip to main content

Search

“All disciplines” · 1449 results

Page 24 · showing 20
arXiv

δ-Vision rebuilds layer-wise visual memory with lightweight MLP adapters, cutting FLOPs to 17.90% while keeping every visual token

Using low-rank interventions, the work finds that visual-to-text information flow in multimodal LLMs is concentrated in a low-dimensional subspace and that layer-specific visual states are highly predictable by lightweight MLPs; it therefore proposes δ-Vision, which freezes the language-model backbone and uses low-rank adapters to construct each layer's visual memory while preserving all visual tokens for text retrieval, reaching an average score of 74.4 on Qwen3-VL-4B (92.9% of uncompressed performance), 9.1 points above the strongest 5%-retention pruning baseline DivPrune, and requiring only 17.90% of the uncompressed model's FLOPs on Video-MME.
arXiv

DISCO splits long context into parallel Worker grounding plus Driver reasoning, holding 78.4% on 1M-token RULER-QA while cutting inference cost by over 80%

The work proposes DISCO, which separates long-context processing into local evidence grounding executed in parallel by lightweight Worker LLMs and global reasoning handled by a central Driver LLM whose dynamic DAG planning is optimized with GRPO reinforcement learning; with Qwen3-8B it holds 78.4% accuracy on RULER-QA at 1M tokens (where the RAG baseline collapses to 10.9%), improves up to 9.8 points over full-context baselines on LongBench v2, and matches Gemini-3-Pro-Preview while reducing inference cost by over 80%.
arXiv

Intervention experiments across DeepSeekMoE, OLMoE and Qwen3-MoE show post-merge expert routing drift is mostly input-representation-induced yet fails to predict task gains from source-route restoration

The work introduces a routing analysis toolkit and, across DeepSeekMoE, OLMoE and Qwen3-MoE, crosses source and merged router inputs and parameters and runs paired token- and task-level evaluations, finding that most expert reassignments are input-shift-induced rather than parameter-induced, that source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and that different expert selections can produce directionally similar mixture outputs; it therefore redefines routing failure as task loss recoverable under a specified routing intervention with non-routing parameters fixed, and proposes Selective Router Repair as a case study that finds no reliable evidence that source-specialist token-likelihood advantages identify beneficial local corr
arXiv

NanoForecast v0.5 fixes only its training pipeline, cutting MASE from 3.030 to 1.704 and beating 200M-parameter TimesFM on all three ETT datasets

Keeping the 6.5M-parameter v0.3 architecture unchanged, NanoForecast v0.5 retrains after fixing loss-scope handling, tensor shape alignment, and augmentation coverage, cutting overall MASE from 3.030 to 1.704 (a 43.8% reduction) under one fixed protocol on the same data and compute budget, and beating 200M-parameter TimesFM on ETTh1, ETTh2, ETTm1, and exchange rate while beating 15M+-parameter PatchTST on all three ETT sets.
arXiv

ColNanoVDR distills multi-vector visual document retrieval queries into a 149M text-only student, keeping about 95% of NDCG@5 and encoding queries 26x faster on one CPU thread

The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
arXiv

Pruned CTC removes linear-in-vocabulary memory growth from large-vocabulary CTC training, and LLM-CTC stays within 7% relative WER of LLM-CE on GigaSpeech while recognizing 7–10× faster

The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
arXiv

ZJU team releases EMem-Bench: across 2,554 long-horizon embodied episodes the strongest model Gemini-3-Flash reaches only 64.2% average success, while its EMem memory system lifts GPT-5.4-mini by 23.6 points

The OmniAI Group of ZJU ACES Lab introduces EMem-Bench, a long-horizon embodied memory benchmark of 2,554 executable episodes spanning Passive Observation, Dynamic Tracking, Interaction Failure, and Experience Generalization, together with an external memory system EMem and an 8B policy EMem-8B; evaluating 16 open-source and proprietary MLLMs plus representative multimodal memory systems shows the strongest proprietary model Gemini-3-Flash reaches only 64.2% average success rate and most open-source models fall below 45%, while EMem achieves the best overall performance among matched backbones, improving Mistral-Small-3.1-24B by 16.4 points and GPT-5.4-mini by 23.6 points, with EMem-8B improving over Qwen3-VL-8B by 20.6 points.
arXiv

CIS recasts training-inference mismatch as a log-odds displacement and tops the five-benchmark average on all three MoE models

The work studies training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and introduces calibrated importance sampling (CIS): the mismatch is characterized as an additive displacement in log-odds, and large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases; across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines, and diagnostic analyses show it places less truncation bias on low-confidence tokens than truncat
arXiv

DCSD decouples credit direction from magnitude in self-distillation, topping 11 benchmarks and lifting math reasoning by 8.45 points over base models

The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
arXiv

Turning the residual stream from passive summation into active retrieval: mirror routing cuts DiT-XL/2 FID from 18.85 to 15.07 and REPA-XL/2 from 5.9 to 4.34

The work systematically analyzes Diffusion Transformer internal representations, finding a latent preference for early-layer feature reuse and symmetric layer pairing, and on that basis proposes a structured connectivity design that turns residual connections from passive summation into active retrieval: the DiT is split evenly along depth into an encoder and a decoder phase, encoder-side representations serve as differentiable residual sources, and each decoder sublayer adaptively fuses its mirrored encoder source through lightweight softmax routing; on ImageNet class-conditional generation, with under 0.1% additional parameters, the method lowers FID from 69.81 to 62.48 on DiT-S/2, from 44.21 to 36.33 on DiT-B/2, and from 18.85 to 15.
arXiv

WideSWE tests coding agents on 120 cross-repository tasks: best configuration solves only 42.50%, and joint execution beats independent runs on requests

Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
arXiv

AdaGuard uses an adaptive guard model to review agent trajectories against user-defined policies, with its 4B model reaching 89.30% binary accuracy on AdaptiveSafety

The work introduces AdaptiveSafety, a dataset of 10,939 training and 1,000 test examples covering policies with 1 to 100 rules, together with SafePO, a reinforcement learning algorithm, and uses them to train the AdaGuard family of 0.6B, 4B, and 8B guard models that assess agent trajectories under policies supplied at inference time; the 4B model achieves binary accuracies of 89.30% on AdaptiveSafety and 71.82% on DynaBench.
arXiv

RoboFoundry treats the agent system itself as the policy to evolve, lifting GPT-5.5 by 27.8% on EmbodiedBench and transferring zero-shot to real robots

RoboFoundry proposes a Self-Evolving System-as-Policy framework that treats the entire supporting system around a foundation model—its context system and hierarchical skill system—rather than model parameters as the policy to evolve, diagnosing capability gaps from execution traces, converting them into validated task-specific system updates, and promoting recurring improvements into general system capabilities; it achieves state-of-the-art results on EmbodiedBench (improving GPT-5.5 by 27.8% and bringing Qwen3.7-Plus to 70.3% versus GPT-5.5's 72.7%), outperforms all baselines on RoboMemArena long-horizon memory by at least 39.0%, beats Cap-Agent0 on LIBERO-PRO by 243.8%–679.7%, and demonstrates zero-shot transfer and online evolution on real robots.
arXiv

Distillation attacks still steal closed-model reasoning after reinforcement learning, with summary traces matching full traces

The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
arXiv

RenderRank reranks documents from compressed visual tokens rendered as images, reaching 55.96 average NDCG@10 on BEIR with 16.5–35.5% fewer input tokens

RenderRank renders document text as images and encodes them into compressed visual tokens, aligning relevance scores to a text-based teacher and then refining positive-versus-negative scores within each query, reaching an average NDCG@10 of 55.96 on 11 BEIR datasets with 290.07 average input tokens, 16.5–35.5% fewer than the evaluated text-based rerankers, and an average NDCG@10 of 88.27 on four long-document datasets at roughly half their average input token count.
arXiv

SCOPD lifts VLM performance at 10% visual-token retention from 86.37% to 92.43% via sparse-context on-policy self-distillation

Using a fixed-context Pass@K analysis, this work shows that degradation after visual-token pruning is not only lost information but also unreliable use of surviving evidence (a representation-utilization gap), and introduces SCOPD and SCOPD+ sparse-context on-policy self-distillation, raising the normalized average over 13 benchmarks at 10% visual-token retention from 86.37% (Vanilla) to 90.49% and 92.43% without architectural changes, extra inference-time computation, or ground-truth answers.
arXiv

Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction

This monograph develops a unified account of training and inference in Power Law Decoder Representation language models (PLDR-LLMs): exact finite work identities decompose changes in the absolute energy of the row-centered learned map into parameter contributions, signed interactions, and numerical observation defects; positive affine blocking retains restarts at the row-constant face while the augmented AdamW state supplies the complete dynamical description; predictive renormalization acts on the complete conditional training law for a single pass over distinct corpus target blocks; experiments reveal observer and optimizer dependence, reject the tested autonomous row-state candidates, and support finite conditional prediction and state-specific operator reduction, while independent sing
arXiv

TokenCast forecasts token consumption during agent runs with composable segment costs, cutting MAE by 14.5% on average across 96 comparisons

TokenCast learns a composable cost triple for each execution segment of an LLM agent (call count, net input-length change, and a cost residual), uses an exact composition identity to fold the cost of earlier context being re-read by every later call into a cumulative estimate, and refreshes the forecast from newly observed execution evidence without additional LLM calls; on SWE-bench Verified the mean cumulative prediction time is 32.8 ms per run, mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations across 4 task suites and 6 agent models, and offline budget-control replay uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion.
arXiv

Imprint Reader uses SMaRT to read frozen weight updates into natural language, and MetaEdit turns those descriptions into behavioral intervention

The work introduces Imprint Reader: Semantic Mount-and-Read Tuning (SMaRT) mounts a frozen weight update onto the Reader and uses an anchor-free meta-query to elicit a natural-language description, with no-change and random-perturbation controls discouraging unsupported claims; on held-out updates the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior, while the Reader's differentiable proxy for a target behavior transfers back to the original model through MetaEdit, raising measured harmful-prompt refusal from 57.9% to 64.1% at a 0.5% pruning rate and lifting BFCL Overall from 41.69% to 44.60% without target-task training data.
arXiv

Moving bidding on device: one tick of sync staleness overspends budgets by 17.77%, and 50 ticks by 1,669.31%

In an auction-logic-faithful on-device simulation with 36 campaigns, 50 devices, and 30 paired demand paths, the study examines two economic misalignments created when privacy-preserving ML decisions move onto clients, finding that proportional Even pacing overspends 17.77% after one tick of staleness and 1,669.31% after 50 ticks, that 50-tick overspend remains 106.95% at two-times budget pressure, that a visible-budget no-sale guard makes zero-lag compliance exact yet leaves 11.88% overspend at one tick, and that 98.23% of rival auctions at one tick admit a profitable deviation once the ML/pacing score transformation is allowed to change payment units.