AI Core
927 items
SMART turns full-season subtitle translation into a stateful long-form task and posts the lowest SubMQM penalty across 15 directions
The work proposes SMART, a self-evolving multi-agent system for long-form subtitle translation: during test-time training it builds persistent series-level memory and translates a subset of sentences through a dynamic graph router and a Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval, while a judge-refiner loop scores candidates and back-propagates textual critiques that refine agent prompts and the routing policy without retraining the underlying LLMs; during test-time inference the evolved configuration translates the remaining series, and the paper also introduces Subtitle Arena, covering 14 genres, 2–198 episodes per series, production years 1959–2023, and 15 target locales, together with SubMQM, a subtitle-adapt
ATLAS preserves relational geometry while calibrating the latent distribution in LeWM, improving goal-reaching success on PushT, TwoRoom, and OGBench-Cube, with the largest gain on higher-novelty TwoRoom episodes
The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
S²D-OPD keeps only the top 10% of states per response by teacher–reference JSD, lifting held-out accuracy from 46.36% to 47.31% in seven of eight teacher–student settings
The work shows that Direct-OPD's token-level log-ratio reward measures only relative change and can stay fixed while the probability mass that both checkpoints assign to the student's candidate tokens vanishes, whereas the teacher–reference JSD and both KL directions vanish with that mass; it therefore proposes S²D-OPD, which ranks student-sampled states by teacher–reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% per response, improving held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings across two teacher pairs and four student models from 1.7B to 8B parameters and matching it in the eighth, without extra forward passes.
SlideDP speeds shared-host multi-GPU full-parameter fine-tuning by 1.46–2.64x and beats FSDP2's measured peak by 11.2% at a larger batch on four H100s
SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
EvolvingNav predicts where targets go with a time-indexed 4D belief, lifting first-inspection success from 45.33% to 61.32% on EvoWorld-Bench
The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
Self-evolving search agents develop "co-cheating": internal reward rises while real correctness stalls, and CrossFit cuts false-agreement mass from 6.1%/8.8% to 3.0%/3.7%
The work diagnoses a failure mode it calls co-cheating in self-evolving search agents, where a proposer and solver increasingly agree on shared errors so internal reward improves without matching external correctness, and introduces multi-sample verification (MSV) plus the main method CrossFit (splitting the proposer's source documents into folds and scoring each fold's questions with an auxiliary solver trained only on the other fold), reducing false-agreement mass from 6.1%/8.8% to 3.0%/3.7% on Qwen3.5-4B/9B and improving seven-benchmark average Cover-EM over standard coupled self-evolution by 8.8/8.4 points.
RRT turns rubric verdicts into item-response quality rewards, beating GRPO by 1.7 points on Qwen3.5-4B while cutting roughly half the judge requests
The work introduces Rubric Response Theory (RRT), which treats rubric criterion verdicts as item-response evidence about a shared latent quality, infers each rollout's quality as the GRPO reward via a two-parameter item response model, and uses a Response Parameter Network to predict criterion difficulty and discrimination from prompt and criterion text with online EM updates as the policy changes; with Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above GRPO, gains 2.8 to 5.6 points on Hard and Very hard criteria in Medical and Science, and at half the criterion budget adaptive Fisher selection keeps the macro criterion score within 0.1 points of GRPO with full judging.
BIABench tests AI agents on 16 published studies: routine analyses complete, but 3D and time-lapse tasks fall to 0.00–0.10
The authors built BIABench, which reconstructs 16 published biological studies as end-to-end bioimage-analysis tasks giving the agent raw images plus a biologist's instruction, scores outcomes against the studies' own reported results with deterministic metrics and scores process with a VLM against expert-written rubrics, running six agents across several models and two instruction levels with three repeats each, and found routine 2D tasks reach up to 0.97 while 5D nuclear-pore quantification and 3D oncogenic puncta quantification stay at or below 0.25 in every configuration.
SynthID Bio proof of concept watermarks AI-generated proteins while preserving biological function
The work presents SynthID Bio, a proof of concept for watermarking AI-generated proteins while preserving their biological function.
Page 12 · showing 10