Skip to main content

AI Core

950 items

  1. arXiv

    RRT turns rubric verdicts into item-response quality rewards, beating GRPO by 1.7 points on Qwen3.5-4B while cutting roughly half the judge requests

    The work introduces Rubric Response Theory (RRT), which treats rubric criterion verdicts as item-response evidence about a shared latent quality, infers each rollout's quality as the GRPO reward via a two-parameter item response model, and uses a Response Parameter Network to predict criterion difficulty and discrimination from prompt and criterion text with online EM updates as the policy changes; with Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above GRPO, gains 2.8 to 5.6 points on Hard and Very hard criteria in Medical and Science, and at half the criterion budget adaptive Fisher selection keeps the macro criterion score within 0.1 points of GRPO with full judging.
  2. arXiv

    BIABench tests AI agents on 16 published studies: routine analyses complete, but 3D and time-lapse tasks fall to 0.00–0.10

    The authors built BIABench, which reconstructs 16 published biological studies as end-to-end bioimage-analysis tasks giving the agent raw images plus a biologist's instruction, scores outcomes against the studies' own reported results with deterministic metrics and scores process with a VLM against expert-written rubrics, running six agents across several models and two instruction levels with three repeats each, and found routine 2D tasks reach up to 0.97 while 5D nuclear-pore quantification and 3D oncogenic puncta quantification stay at or below 0.25 in every configuration.
  3. Google DeepMind News

    SynthID Bio proof of concept watermarks AI-generated proteins while preserving biological function

    The work presents SynthID Bio, a proof of concept for watermarking AI-generated proteins while preserving their biological function.
  4. MIT News - Artificial intelligence

    MIT-led team's Ataraxos beats the world's strongest Stratego player 15-1-4, training on under one hundredth of DeepNash's examples

    Researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI system called Ataraxos that combines a self-play reinforcement learning "blueprint strategy" with decision-time planning using a generative model, defeating the strongest human Stratego player by a record 15-1-4 margin and achieving a 39-2 record against top players at the Stratego world championship, while using less than one hundredth of DeepMind's DeepNash training examples and less than one thirtieth of its self-play games, and generalizing to other imperfect-information games such as Barrage Stratego, Hanabi, and Dou dizhu.
  5. Cohere Labs

    Cohere proposes RCP-nDCG@10: a calibrated AI judge replaces fixed answer keys, matching human preference in 77% of 289 blind contests versus 52% for conventional nDCG

    Cohere introduces RCP-nDCG@10, a retrieval evaluation methodology in which a calibrated AI judge grades every retrieved document against the same explicit relevance rubric and combines rubric answers with group-wise comparisons and calibration, so relevant documents absent from the answer key can still earn credit; in a blind study with 46 annotators, 273 queries and 289 head-to-head contests, reviewers sided with RCP-nDCG 70% of the time where the two metrics named different winners, and across all contests RCP-nDCG matched reviewer preference 77% of the time versus 52% for conventional nDCG, with the gain driven mainly by crediting relevant documents the answer key missed (19 percentage points).
  6. Cohere Labs

    Cohere launches Embed 5 with Pro at 85.8 average on ViDoRe V3 and Fast at 2.4x throughput sharing one embedding space

    Cohere released the Embed 5 family of embedding models in two tiers, Pro and Fast, which share a single embedding space and support 128K context, text/image/fused inputs, 100+ languages, and Matryoshka outputs from 2048 down to 256 dimensions; Pro averages 85.8 on ViDoRe V3 (an 8.8 gain over Embed 4), scores 80.1 on FinanceBench and 84.8 across the parsed-document suite, while Fast averages 84.5 on ViDoRe V3 with about 2.4x the document throughput of Pro, priced at $0.12 and $0.08 per million tokens respectively.
  7. Journal of Computer Science and Technology

    Survey maps high-level synthesis for approximate computing around error estimation, approximation techniques, and design space exploration, and flags research gaps

    Addressing the lack of a systematic survey and in-depth analysis of the latest methodologies in high-level synthesis for approximate computing (AHLS), this survey summarizes recent technologies in the field with particular focus on error estimation, approximation techniques, and design space exploration (DSE), and analyzes current research gaps, aiming to give researchers, engineers, and scholars a theoretical and practical framework for AHLS.
  8. Natural Sciences and Applied Technology

    FedTrust-GNN reaches 94.2% accuracy at 10,000-100,000 participants and cuts label-flipping attack success from 34% to 6.1%

    The work proposes FedTrust-GNN, a decentralized user-modeling framework that combines differentially private federated learning with secure multi-party computation, a permissioned blockchain using PBFT consensus, and a heterogeneous graph attention network (HGAT) that infers dynamic trust scores, with trust-weighted robust aggregation (TWRA, combining norm clipping and coordinate-wise median aggregation) providing Byzantine fault tolerance; on Federated EMNIST, Stack Overflow, and synthetic datasets with 10,000-100,000 participants it reports 94.2% accuracy (within 1.3% of centralized models), a reduction of label-flipping attack success from 34% to 6.1% (an 82% reduction), a 41% improvement in convergence stability, and blockchain performance of 1,200 TPS with 2.3-second finality.
  9. Natural Sciences and Applied Technology

    DCONVNET splits 2D direction-of-arrival estimation into two 1D problems and estimates azimuth and elevation via the alternating direction method of multipliers

    The work presents a fast two-dimensional direction-of-arrival (DOA) estimation approach for low-elevation targets of very-high-frequency array radar: it uses the azimuth and pitch angle uncoupling properties of a uniform planar array to turn the 2D angle estimation problem into two 1D DOA estimation problems, retrieves target information in the azimuth and elevation dimensions with digital beamforming, and then estimates azimuth and pitch angles using the alternating direction method of multipliers, thereby reducing complexity and eliminating the need for eigenvalue decomposition during operation.
  10. medRxiv

    1,003 people with depression rated psychotherapy with different levels of AI involvement: they preferred human therapists, willing to pay 31.6% less for assistive or collaborative AI and 57.2% less for fully autonomous AI

    The study had 1,003 participants with depression read vignettes describing psychotherapy options with different levels of AI involvement (a human therapist without AI, assistive AI, collaborative AI, and fully autonomous AI) and rate them; participants consistently evaluated human therapists more favorably, reporting greater likelihood of seeking treatment, less hesitancy, and greater treatment acceptability, and compared with a human therapist they were willing to pay 31.6% less for therapists using assistive or collaborative AI and 57.

Page 15 · showing 10