Skip to main content

AI Core

927 items

  1. arXiv

    Jev holds high NDCG@10 across three Amazon domains while its latency grows far more gradually than pointwise Qwen

    On Movies and TV, Video Games, and Books from Amazon Reviews 2023, the study builds controlled hard candidate sets from SASRec's highly ranked negatives (candidate sizes 10 to 100) and compares Jev, a decision-oriented "System One Model" from TypeSafe AI, against SASRec, DCNv2, and pointwise and listwise Qwen2.5 7B Instruct rerankers, finding that Jev maintains strong recommendation effectiveness (NDCG@10, Hit Rate@10, MRR) while its observed serving latency grows substantially more gradually than pointwise Qwen, though it remains far above recommendation-specific models.
  2. arXiv

    SMART turns full-season subtitle translation into a stateful long-form task and posts the lowest SubMQM penalty across 15 directions

    The work proposes SMART, a self-evolving multi-agent system for long-form subtitle translation: during test-time training it builds persistent series-level memory and translates a subset of sentences through a dynamic graph router and a Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval, while a judge-refiner loop scores candidates and back-propagates textual critiques that refine agent prompts and the routing policy without retraining the underlying LLMs; during test-time inference the evolved configuration translates the remaining series, and the paper also introduces Subtitle Arena, covering 14 genres, 2–198 episodes per series, production years 1959–2023, and 15 target locales, together with SubMQM, a subtitle-adapt
  3. arXiv

    ATLAS preserves relational geometry while calibrating the latent distribution in LeWM, improving goal-reaching success on PushT, TwoRoom, and OGBench-Cube, with the largest gain on higher-novelty TwoRoom episodes

    The authors introduce ATLAS, a training objective that transfers normalized pairwise structure from an encoder's mean-pooled patch features to the planning latent and uses Wasserstein embedding matching (WEMReg) based on one-dimensional Wasserstein-2 transport to calibrate its marginal; instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes, while diagnostics show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error.
  4. arXiv

    S²D-OPD keeps only the top 10% of states per response by teacher–reference JSD, lifting held-out accuracy from 46.36% to 47.31% in seven of eight teacher–student settings

    The work shows that Direct-OPD's token-level log-ratio reward measures only relative change and can stay fixed while the probability mass that both checkpoints assign to the student's candidate tokens vanishes, whereas the teacher–reference JSD and both KL directions vanish with that mass; it therefore proposes S²D-OPD, which ranks student-sampled states by teacher–reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% per response, improving held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings across two teacher pairs and four student models from 1.7B to 8B parameters and matching it in the eighth, without extra forward passes.
  5. arXiv

    SlideDP speeds shared-host multi-GPU full-parameter fine-tuning by 1.46–2.64x and beats FSDP2's measured peak by 11.2% at a larger batch on four H100s

    SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
  6. arXiv

    EvolvingNav predicts where targets go with a time-indexed 4D belief, lifting first-inspection success from 45.33% to 61.32% on EvoWorld-Bench

    The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
  7. arXiv

    Self-evolving search agents develop "co-cheating": internal reward rises while real correctness stalls, and CrossFit cuts false-agreement mass from 6.1%/8.8% to 3.0%/3.7%

    The work diagnoses a failure mode it calls co-cheating in self-evolving search agents, where a proposer and solver increasingly agree on shared errors so internal reward improves without matching external correctness, and introduces multi-sample verification (MSV) plus the main method CrossFit (splitting the proposer's source documents into folds and scoring each fold's questions with an auxiliary solver trained only on the other fold), reducing false-agreement mass from 6.1%/8.8% to 3.0%/3.7% on Qwen3.5-4B/9B and improving seven-benchmark average Cover-EM over standard coupled self-evolution by 8.8/8.4 points.
  8. arXiv

    RRT turns rubric verdicts into item-response quality rewards, beating GRPO by 1.7 points on Qwen3.5-4B while cutting roughly half the judge requests

    The work introduces Rubric Response Theory (RRT), which treats rubric criterion verdicts as item-response evidence about a shared latent quality, infers each rollout's quality as the GRPO reward via a two-parameter item response model, and uses a Response Parameter Network to predict criterion difficulty and discrimination from prompt and criterion text with online EM updates as the policy changes; with Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above GRPO, gains 2.8 to 5.6 points on Hard and Very hard criteria in Medical and Science, and at half the criterion budget adaptive Fisher selection keeps the macro criterion score within 0.1 points of GRPO with full judging.
  9. arXiv

    BIABench tests AI agents on 16 published studies: routine analyses complete, but 3D and time-lapse tasks fall to 0.00–0.10

    The authors built BIABench, which reconstructs 16 published biological studies as end-to-end bioimage-analysis tasks giving the agent raw images plus a biologist's instruction, scores outcomes against the studies' own reported results with deterministic metrics and scores process with a VLM against expert-written rubrics, running six agents across several models and two instruction levels with three repeats each, and found routine 2D tasks reach up to 0.97 while 5D nuclear-pore quantification and 3D oncogenic puncta quantification stay at or below 0.25 in every configuration.
  10. Google DeepMind News

    SynthID Bio proof of concept watermarks AI-generated proteins while preserving biological function

    The work presents SynthID Bio, a proof of concept for watermarking AI-generated proteins while preserving their biological function.

Page 12 · showing 10