Skip to main content

AI Core

998 items

  1. arXiv

    WideSWE tests coding agents on 120 cross-repository tasks: best configuration solves only 42.50%, and joint execution beats independent runs on requests

    Drawing on 1,729,171 pull requests across 103 software ecosystems, the authors built WideSWE, 120 real tasks (60 bug fixes, 60 features) that require coordinated changes across multiple repositories, and evaluated seven coding-agent configurations: full task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest, while joint and per-repository independent execution each help in different ways.
  2. arXiv

    DN-MOPD normalizes per-domain teacher feedback and lifts the six-task average over MOPD from 58.4/50.3/26.6 to 59.6/52.5/29.0 across three Qwen3.5 sizes

    The work finds that multi-teacher on-policy distillation (MOPD) with domain-label routing alone fails to beat the strongest single-teacher student and transfers little of the mathematics expert's gain, because domain feedback is unbalanced (in the first training batch instruction-following log-ratios are 2.3-4.4 times as dispersed as the pooled signal while mathematics is about half as dispersed, and for the initial 4B student the instruction-following loss supplies 94% of the combined gradient), and proposes DN-MOPD, which keeps label routing and rescales each domain's distillation advantages by the clipped ratio of pooled to domain log-ratio spread estimated every batch, improving the six-task average over MOPD at every size and both evaluation budgets (1.17-2.36 points at 16K, 2.47-3.
  3. arXiv

    ColNanoVDR distills multi-vector visual document retrieval queries into a 149M text-only student, keeping about 95% of NDCG@5 and encoding queries 26x faster on one CPU thread

    The work introduces ColNanoVDR, which uses OTW, an entropic optimal-transport objective with learned token weights, to distill the query encoder of multi-vector visual document retrieval teachers into 149M text-only students without encoding or reading any page during training, retaining about 95% of the teachers' NDCG@5 on ViDoRe v1-v3, encoding queries up to 26x faster on a single CPU thread, and matching score distillation under identical training while reading 12.6x less cached teacher data.
  4. arXiv

    DCSD decouples credit direction from magnitude in self-distillation, topping 11 benchmarks and lifting math reasoning by 8.45 points over base models

    The work introduces Decoupled Credit Self-Distillation (DCSD), which uses belief-margin probing to set each reasoning step's credit direction and marginal information gain to quantify its contribution magnitude, leaving the privileged teacher to allocate credit only within steps; across 11 mathematical and multimodal reasoning benchmarks DCSD achieves the best overall scores against GRPO, OPSD, RLSD and RLCSD, improving overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning over base models, while correcting credit direction for about 6% of tokens and reducing token credit magnitude by roughly 1.5 times.
  5. arXiv

    Swapping matrix multiplication for an associative-algebra product: a 110M-parameter model gains 6.2–7.8% generation throughput while GSM8K, MBPP and IFEval all drop

    The work replaces ordinary matrix multiplication in Transformer projections with an associative-algebra multiplication table that keeps the full learned weight bank and parameter count, lowers the bilinear rank for the q=2 case from 7 (Strassen's 2×2 algorithm) to 6, and in a controlled pretraining run of two approximately 110M-parameter models over 12.3B tokens observes a 6.2–7.8% end-to-end generation throughput gain across four prompt domains together with lower scores on GSM8K, IFEval and MBPP than the dense baseline.
  6. arXiv

    Distillation attacks still steal closed-model reasoning after reinforcement learning, with summary traces matching full traces

    The work proposes a more realistic threat model for distillation attacks in which attackers continue training with reinforcement learning after distillation, and shows experimentally that distillation followed by RL outperforms either alone, that a well-known defense (antidistillation sampling) can appear effective after distillation but be broken after RL, and that distilling on approximate reasoning traces reconstructed from closed-source API summaries and final answers, followed by RL, yields reasoning gains similar to using full hidden traces.
  7. arXiv

    Nereus adapts RL post-training parallelism on the fly: 27.7% lower average step latency on a real-data trace and up to 7.27x OpenRLHF throughput for 8B PPO

    Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
  8. arXiv

    Pruned CTC removes linear-in-vocabulary memory growth from large-vocabulary CTC training, and LLM-CTC stays within 7% relative WER of LLM-CE on GigaSpeech while recognizing 7–10× faster

    The work introduces Pruned CTC, which exploits the fact that every valid CTC alignment uses only target tokens and blank to restrict alignment computation to that batch-level subset while retaining full-vocabulary normalization, and proves this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and first-order gradients; with a Zipformer-M encoder and a 180K vocabulary it cuts full-step memory by 5.1× at only 17% step-time overhead and matches standard CTC accuracy across three corpora; built on it, LLM-CTC adapts pretrained causal LLMs to non-autoregressive ASR over native vocabularies, staying within 7% relative WER of LLM-CE across six Qwen3 sizes from 0.
  9. arXiv

    YuE2 writes a readable score before rendering, and experts prefer symbolic planning 49.3% to 34.6% on overall quality

    YuE2 uses a single roughly 3.58B-parameter AR–NAR Mixture-of-Transformers to first write an ABC score specifying melody, chords, key, meter, tempo, and form, then expand it into semantic music tokens and render full-song audio, with companion models MERT2 and SheetSage2 constructing semantic and symbolic supervision from recordings; with the same checkpoint, experts prefer symbolic planning 49.3% to 34.6% for overall quality, YuE2 scores 6.7316 on SongBench Global Avg on WildSongBench (6.9632 with best-of-8), MERT2 surpasses previous best results on 14 of 15 MARBLE metrics, and SheetSage2-AR leads 12 of 15 benchmark–metric pairs.
  10. arXiv

    GEAR treats geometry as a memory address: sparse correspondence attention cuts ATE by 80.3% for minute-long camera-controlled video generation

    The work introduces GEAR, which avoids fusing history into a persistent global 3D representation and instead uses per-frame geometry to build token-level history-to-target correspondences, treating geometry as an explicit address for visual memory; Geometric Correspondence Attention then sparsely injects only geometrically matched historical features, while an Invisible Octree rejects projectable but occluded correspondences, reporting state-of-the-art visual quality, camera control, and revisit consistency on DL3DV-Evaluation and WorldScore-Static and enabling minute-long generation.

Page 33 · showing 10