Skip to main content

AI Core

1004 items

  1. arXiv

    DepthBench sweeps width–depth ratio across 10 architectures at fixed parameter budget: HC and Full AttnRes keep lowering validation loss at extreme deep-narrow shapes while Pre-LN degrades

    The work introduces DepthBench, a controlled benchmark that varies the width–depth aspect ratio (d_model/n_layer) from shallow-wide to deep-narrow under a fixed parameter budget and pre-training recipe, comparing 10 representative residual and normalization architectures, and finds that the benefit of depth is strongly architecture-dependent: Pre-LN and most of its norm- and scaling-based variants worsen as models get deeper and narrower (Pre-LN rises from 2.759 to 2.782), whereas HC and Full AttnRes improve consistently (HC from 2.729 to 2.699, Full AttnRes from 2.751 to 2.718), accompanied by stronger cross-layer causal dependence and order sensitivity, though deeper models also incur compute and memory overhead.
  2. arXiv

    Xiaomi's groupwise agentic grading and advantage redistribution lifts 310B and 1.02T code agents to 67.9 and 71.9 on DeepSWE

    Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
  3. arXiv

    Skill2Env synthesizes 2,963 executable environments from skills, and 1.5K fine-tuning trajectories lift a 35B agent from 36.6 to 45.0 across seven benchmarks

    Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
  4. arXiv

    SkillDRE evolves malicious agent skills through a dual-stage loop of pre-execution scanning and runtime feedback, reaching a 45.28% average attack success rate across four victim models on SkillsBench

    SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
  5. arXiv

    PReCache shares KV caches across multi-LoRA agents via low-rank precomputation and neutral reconstruction, cutting TTFT by up to several times with almost no accuracy loss

    The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
  6. arXiv

    Duplex-MPE tests 2,000 multi-party dialogue scenarios and finds that talking more is not answering better, with MiniCPM-o 4.5 leading three of four capabilities

    The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
  7. arXiv

    EditWorld moves video world models from navigation to streaming editing, scoring 73.8 overall and 80.0 on editing in WBench-Editing

    EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
  8. arXiv

    Relic turns recurring collaboration failures into executable protocols, lifting complete-contract delivery from 14.06% to 19.76% across 360 controlled runs

    Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
  9. arXiv

    NUS team introduces a residual-transferability metric and CoverLock, cutting forgery success below 0.25% across three high-RT watermarking systems

    The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
  10. arXiv

    AdaTutoRank turns a set-level rubric score into token-level credit via adaptive tutoring, reaching the best overall score across ten benchmarks with fewer retrieval calls

    The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.

Page 36 · showing 10