AI Core
1004 items
Xiaomi's groupwise agentic grading and advantage redistribution lifts 310B and 1.02T code agents to 67.9 and 71.9 on DeepSWE
Xiaomi introduces GAGAR, a framework for code agent RL in which an SFT-trained agentic grader jointly inspects test-passing trajectories within a task group, ranks them along five quality dimensions, and redistributes advantage from lower- to higher-quality solutions in a sum-preserving way; in controlled code-only experiments with MiMo-V2.6-Flash (310B total parameters) it reaches a 62.2% DeepSWE v1.1 pass rate at step 28, 12.1 percentage points above the binary-reward baseline, and after industrial-scale mixed-task RL with Flash and Pro (1.02T total parameters) the checkpoints reach DeepSWE v1.1 avg@3 of 67.9 and 71.9.
Skill2Env synthesizes 2,963 executable environments from skills, and 1.5K fine-tuning trajectories lift a 35B agent from 36.6 to 45.0 across seven benchmarks
Skill2Env is a capability-oriented framework that curates skills from public skill libraries, uses reusable difficulty patterns and task blueprints to jointly synthesize task instructions, execution substrates, workspaces, and rubric-based evaluators, and iteratively hardens tasks using solver execution evidence, yielding 2,963 executable tasks; supervised fine-tuning of Qwen3.6-35B-A3B on 1.5K high-scoring trajectories from these environments raises the unweighted average across seven agent benchmarks from 36.6 to 45.0.
SkillDRE evolves malicious agent skills through a dual-stage loop of pre-execution scanning and runtime feedback, reaching a 45.28% average attack success rate across four victim models on SkillsBench
SkillDRE is a fully automated framework that, given a benign task and its associated skills, autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule, then holds both fixed while alternating scanner-guided evolution with runtime-defense-guided refinement, forming a closed loop between the pre-execution and runtime stages; on SkillsBench across four victim models it achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance.
PReCache shares KV caches across multi-LoRA agents via low-rank precomputation and neutral reconstruction, cutting TTFT by up to several times with almost no accuracy loss
The work presents PReCache, a training-free KV cache sharing framework with two designs: PreLRShared precomputes a compact agent-specific low-rank cache for every agent when each segment of the shared context is first processed, removing later agents' repeated prefill over the accumulated trajectory, while ReBaseShared further reconstructs the shared base cache from adapter-free hidden states to reduce dependence on the previous agent, with lazy prefill for single-stream inference and double batching for concurrent serving; across LLaMA-3.
Duplex-MPE tests 2,000 multi-party dialogue scenarios and finds that talking more is not answering better, with MiniCPM-o 4.5 leading three of four capabilities
The authors introduce Duplex-MPE, a benchmark of 2,000 paired scenarios with three or four human speakers plus an assistant named Aria, contrasting explicit naming with implicit addressing, and evaluate five open-weight full-duplex speech systems on continuous audio without transcripts, speaker labels or turn boundaries using four scores—fresh response initiation, conditional answer accuracy, silence preservation and answering-window yield—finding that MiniCPM-o 4.5 leads on three scored capabilities while frequent speech from other systems often coexists with inaccurate answers or failures to remain silent.
EditWorld moves video world models from navigation to streaming editing, scoring 73.8 overall and 80.0 on editing in WBench-Editing
EditWorld is a video world model that accepts streaming editing instructions and reference images during autoregressive generation through Gated Causal Attention, a Sparse Context mechanism, joint autoregressive and bidirectional training, and a dedicated data synthesis and annotation pipeline, and it introduces WBench-Editing with roughly 150 cases of 240–480 frames each, achieving the best overall score of 73.8 and an editing score of 80.0 on that benchmark.
Relic turns recurring collaboration failures into executable protocols, lifting complete-contract delivery from 14.06% to 19.76% across 360 controlled runs
Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.
NUS team introduces a residual-transferability metric and CoverLock, cutting forgery success below 0.25% across three high-RT watermarking systems
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
AdaTutoRank turns a set-level rubric score into token-level credit via adaptive tutoring, reaching the best overall score across ten benchmarks with fewer retrieval calls
The work proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level, nine-dimension rubric hierarchy that assigns each rollout one of three hint forms—rubrics, a corrective reflection, or a sibling set—matched to its reward, and distills the hint's effect into a token-level advantage combined with the group-relative outcome advantage; across ten benchmarks spanning RAG, deep research, and setwise evaluation it attains the best overall performance while reducing the agent's retrieval calls.
Page 36 · showing 10