Skip to main content

AI Core

915 items

  1. arXiv

    First joint scaling law for looped mixture of experts: sparsity raises and delays the saturation of recurrence's effective-parameter gain

    The work introduces Loop Scaling Laws, the first scaling law that jointly models recurrence and MoE sparsity alongside model size and training data, using a bounded, sparsity-conditional recurrence mapping to characterize the effective-parameter gain from looping and its saturation; it predicts held-out loss more accurately than linear and power-law mappings, guides the choice of recurrence and expert count under compute and memory budgets, and at trillion-token scale lets a 0.3B-active/1.3B-total looped MoE match a roughly 2x larger non-looped MoE on reasoning benchmarks.
  2. arXiv

    DyRAD renders full range-azimuth-Doppler radar tensors of dynamic driving scenes through a fixed point-spread function, raising radar detection recovery on RADIal from 26.9% to 90.7% of reference-detected objects

    DyRAD represents dynamic driving scenes as static background point reflectors plus motion-tracked dynamic point reflectors and renders complete range-azimuth-Doppler (RAD) tensors through a fixed analytic point-spread function derived from the radar signal-processing chain, with Doppler serving both as a rendered output and as supervision for object tracks, evaluated on RADIal, Boreas, and a synthetic benchmark at both on-path poses and displaced viewpoints, raising radar detection recovery on RADIal from 26.9% for the strongest baseline to 90.7% of reference-detected objects.
  3. arXiv

    Imagine3D-LLM has multimodal LLMs imagine a compact 3D scene before answering, reaching 68.5 average on SPAR-Bench at 7B scale

    The work introduces Imagine3D-LLM: a small set of learnable Gaussian summary tokens is appended after the image tokens, decoded into a compact 3D Gaussian Splatting representation, and trained jointly with a photometric reconstruction loss and standard next-token prediction, so that a multi-view multimodal LLM assembles a coarse 3D scene representation before answering; it consistently outperforms prior approaches across seven spatial reasoning and 3D scene understanding benchmarks, reaching 68.5 average on SPAR-Bench and surpassing the strongest spatially-aware baseline 3DThinker-7B by 5.2 points.
  4. arXiv

    RWTD lifts a one-step generator's GenEval from 0.73 to 0.80 via reward-weighted transport distillation

    The work introduces Reward-Weighted Transport Distillation (RWTD), a post-training method for one-step generators that needs only generated samples and scalar reward evaluations, builds an adaptive target mixing separately reward-tilted current and reference distributions, realizes it through feature-space optimal transport and fixed-point regression, and raises the one-step SANA Sprint 1.6B GenEval score from 0.73 to 0.80 while showing cross-reward generalization in preference alignment experiments.
  5. arXiv

    Tacit-TTS swaps autoregressive decoding for masked prediction to make zero-shot voice cloning over 10x faster, staying usable with eight cross-lingual references plus infant babble and synthetic gibberish

    Tacit-TTS, distilled from IndexTTS2, replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, adds training-free acoustic length estimation, and accelerates the flow-matching renderer via ReFlow distillation, achieving competitive zero-shot cloning quality on two English and two Mandarin datasets while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds and supporting cross-lingual and non-lexical references without a reference transcript.
  6. arXiv

    OSWorld-Science tests 12 VLMs on 146 scientific-software tasks: the strongest agent averages 73.7% and only 40% on the hardest set

    The work introduces OSWorld-Science, a benchmark and evaluation environment combining expert-defined scientific tasks, artifact-based execution graders, and a shared agent harness, and evaluates 12 vision-language models across 7 scientific domains, 17 software packages, and 3 languages, finding that the strongest model, Claude Fable 5.1, averages 73.7% and about 40% on the hardest task set, with failures concentrated in the interaction layer and result verification rather than in scientific reasoning itself.
  7. arXiv

    RL post-training of vision-language-action models concentrates parameter updates in Timestep Modules that hold only 14.83%–27.58% of the action expert, and their low-rank directions predict and improve task success

    This work systematically analyzes how reinforcement learning (RL) reshapes flow-based vision-language-action (VLA) models, finding that across π0.5 and GR00T N1.5/N1.6 on LIBERO, ManiSkill, MetaWorld, and CALVIN, RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules, which hold 14.83%–27.58% of the parameters yet capture a disproportionate share of the RL gain, with shift-vector update directions predicting task success at up to 99.6% ROC-AUC and steering along them improving policies without additional RL training.
  8. arXiv

    Selecting distillation positions by entropy shift lifts LLM judges 2–9 points over Dr. GRPO on subjective tasks

    This work studies training LLM judges from natural language feedback, proposes using the per-position entropy shift between teacher and student to distinguish context sharpening from context spreading, and masks high-entropy-shift positions so that self-distilled judges outperform outcome-supervised RL (Dr. GRPO) by 2–9 percentage points on the evaluated subjective subcategories while remaining competitive on objective ones.
  9. arXiv

    GGSD turns five buttons into playable skills via 1v1 self-play: humans clear Maze and CubePush on Ant, Franka and G1 with no extra training

    The work presents Game-Guided Skill Discovery (GGSD), in which a hierarchical agent self-plays 1v1 competitive games against a pool of past checkpoints, a high-level policy selects among 5 discrete skills (6 for G1) and a skill-conditioned low-level policy outputs motor actions, with a mutual-information reward separating skill semantics; after training a human can replace the high-level policy and use the same discrete skills, composing them on Ant, a Franka arm and a Unitree G1 to solve unseen Maze and CubePush tasks, with human success rates of at least 84%.
  10. arXiv

    AED turns 50,228 failed agent rollouts into error–diagnosis pairs: first-proposal corrections lift verifier pass rates from 18.4% to 51.1%, and full-diagnosis fine-tuning raises Qwen3-8B's exact-step agreement from 47.2% to 63.6%

    The work introduces the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families and 23 policy models, and a five-stage Agentic Error-to-Training (AET) pipeline that links natural failures to diagnoses, proposed corrections and optional replay evidence; across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1% (a 32.7 percentage-point gain), full-diagnosis fine-tuning on a separately frozen diagnosis release raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6% (averaged over three seeds on a 943-case holdout), and in a single-seed actor-training comparison action-only repair training scores 6.

Page 9 · showing 10