Skip to main content

AI Core

1001 items

  1. arXiv

    NTU team finds speaker-verification EER does not track human voice similarity, and an embedding-dimensionality bottleneck lifts alignment correlation from 0.08 to 0.74

    Using 16,905 different-speaker pairs from VoxSim, the study builds a human perceptual alignment metric and systematically compares seven training objectives across five model conditions, finding that verification accuracy (EER) does not track human similarity judgments while embedding effective dimensionality correlates with perceptual alignment at about -0.95 rank correlation, and that a dimensionality bottleneck on ECAPA-TDNN raises AM-Softmax alignment from 0.08 to 0.74.
  2. arXiv

    Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs

    Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.
  3. arXiv

    VGGT-Diff routes VGGT-Ω geometry latents into Wan video diffusion, topping PSNR, LPIPS and DreamSim on 6,188 DL3DV targets

    VGGT-Diff is a geometry-routed multi-view diffusion framework that uses a frozen VGGT-Ω to transform source-view features together with 3D points and confidence into query-aligned conditions via a confidence-aware Visual Geometry Router, injects them into a pretrained Wan2.1-I2V-14B video diffusion model, and regularizes predicted-clean residuals along reliable 3D tracks with Point-Track Residual Consistency; trained on only 980 DL3DV scenes, it ranks first in PSNR, LPIPS and DreamSim and second in SSIM on 6,188 DL3DV-Benchmark targets, and first in LPIPS and DreamSim with second-place PSNR on zero-shot Mip-NeRF 360.
  4. arXiv

    QwenGyre pairs elastic GPU scheduling with trajectory-tree processing to lift Qwen 3.8 2.4T from 52.5% to 58.5% on NL2RepoBench in 48 steps

    QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
  5. arXiv

    MT-OPSD uses on-policy self-distillation on self-generated states to lift open-source image editors from 0.03–0.15 to 0.38–0.52 ten-turn success

    The work finds that instruction-based image editing models degrade rapidly under recursive multi-turn editing and attributes this to a train–test mismatch in the conditioning distribution, since models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference; it proposes MT-OPSD, which rolls out the current student under an identity instruction to produce self-generated states containing its own errors, then uses the same model conditioned on the clean source as a teacher to transfer clean-condition editing behavior onto those states via sparse velocity matching, without multi-turn annotations; it also builds LME-Bench of 100 ten-turn editing sessions, raising SR@10 to 0.38–0.52 and cutting CR@10 to at most 0.
  6. arXiv

    DRM replaces the scalar reward head with a diffusion head: matching or beating same-data, same-backbone baselines on five benchmarks while its reward distribution turns multimodal as human disagreement grows

    The work introduces DRM (Diffusion Reward Model), which recasts reward modeling as conditional density estimation: on top of a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family assumed, a single architecture handles both multi-attribute regression and pairwise preference data, it matches or surpasses the same-data, same-backbone scalar head ArmoRM and quantile head QRM across five benchmarks, and it exhibits multimodal structure tied to human disagreement on repeated-annotation data, with uncertainty-aware rejection, LCB aggregation, and downstream RLHF experiments demonstrating the practical value of distributional information.
  7. arXiv

    AI Night-Scientist uses reinforcement learning to teach models when to depart from predictable reasoning, widening research-direction coverage by 27.8% and contribution types by 14.9%

    The work introduces AI Night-Scientist, a cognitive-science-grounded agentic framework that represents creativity at the action, process, and outcome levels and uses GRPO to train Qwen3-8B/14B models for long-horizon research proposal generation, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model, improving predicted citation impact by up to 32.0 percentage points and originality by 66.2 points, and showing that these gains cannot be reproduced by raising decoding temperature alone, with semantic guidance about what kind of creativity to pursue being critical.
  8. arXiv

    DepthBench sweeps width–depth ratio across 10 architectures at fixed parameter budget: HC and Full AttnRes keep lowering validation loss at extreme deep-narrow shapes while Pre-LN degrades

    The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
  9. arXiv

    Xiaomi's groupwise agentic grading and advantage redistribution lifts 310B and 1.02T code agents to 67.9 and 71.9 on DeepSWE

    Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
  10. arXiv

    Skill2Env synthesizes 2,963 executable environments from skills, and 1.5K fine-tuning trajectories lift a 35B agent from 36.6 to 45.0 across seven benchmarks

    Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.

Page 35 · showing 10