Skip to main content

Daily report

AI and science frontiers · 2026-09-16

Only content delivered through the publication boundary on this date is included.

arXiv

Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN

This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
arXiv

EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation

EviGen proposes a three-layer framework that first retrieves and ranks evidence from a patient's longitudinal EHR by predictive contribution rather than textual relevance using learnable queries trained on outcome labels, then has an LLM generate a citation-grounded clinical rationale within that evidence scaffold, and finally applies a process-supervised verifier to check each reasoning step and flag unreliable claims; across MIMIC-IV one-year mortality, autism, and ADHD tasks it improves prediction performance and rationale faithfulness over full-context and RAG baselines and is preferred by most reviewers in a clinical pilot.
arXiv

Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

The work proposes that running tool calls explicitly report their own progress (a fraction of work remaining or an accurate signal that the end is near), and through a census of four public agent corpora, a harness that recovers the signal without changing what the agent sees, a comparison against four published predictors, and an end-to-end evaluation plugged into vLLM through a few small hints, finds that the reported progress is between several times and an order of magnitude more accurate than the best published predictors at KV cache decision points and stays accurate when the environment changes, cutting p90 time to first token after a tool call by 20.7% (HBM only) and 20.8% (HBM + DRAM) against LRU, close to an oracle.
arXiv

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

The work proposes an "infinite-parameter LLM" architecture in which a compact hypernetwork encodes run-time data (facts, instructions, demonstrations) into a low-dimensional latent code, that code generates a low-rank additive modulation of a shared base feed-forward network so each token's "expert" is generated rather than drawn from a stored bank, and a Bayesian belief over the latent code is carried and updated online so the effective weight keeps evolving within a session; the authors also specify an evaluation protocol that pits carrying knowledge in the weights against carrying it in the prompt at matched budget.
arXiv

Degree-Free Spectral Independence for Log-Concave Holant Measures

The paper establishes a degree-independent spectral independence bound for log-concave Holant problems on simple graphs, yielding relaxation-time bounds for Glauber dynamics of O_λ(m) for the monomer–dimer model at activity λ, O_{b,λ}(m) for b-matchings at fugacity λ>0, and O(bm) for uniform b-matchings, where m is the number of edges.
arXiv

Using OCR Heads to Verbalize Image Semantics: General Verbalization Heads and a Verbalization Lens in VLMs

Across four vision-language models (Qwen3-VL-2B, Qwen3-VL-8B, Molmo2-7B, Llava-Next-34B), the work identifies attention heads causally necessary for OCR and shows they are general-purpose verbalization heads; summing their output-value matrices yields a verbalization lens that decodes image-token hidden states into interpretable semantic words at any layer, including layer 0, and whose pseudo-inverse provides concept vectors for editing image representations (e.g., replacing a tractor with a revolver), supporting the view that image representations are aligned with language space from early layers.
arXiv

Compositional Policy Violations: When Step-Level Compliance Fails in Agentic AI Workflows

The paper defines and formalizes Compositional Policy Violations (CPVs), a governance failure mode in which every step of an agentic workflow passes its own local check while the composed execution violates the governing policy, and it offers a four-part taxonomy (Authority Creep, Threshold Laundering, Cumulative Sum Violation, Context Collapse), argues that the correct repair topology is dictated by where the guarded quantity mutates, and proposes a provenance-aware runtime architecture that evaluates policies over complete execution traces and recomputes guarded quantities from raw provenance rather than the pipeline's derived representation.
Terence Tao blog RSS

Proofs, Prompts and Posts: A Community Blog on AI in Mathematics

This guest post by the editors of the communal blog Proofs and Prompts describes why the blog was started and what it aims to do: as AI became a central topic of conversation in the mathematical community, the editors created a collective space for mathematicians who lack a natural platform, especially PhD students, and reports that about a month after launch contributors ranged from PhD students and undergrads to Fields medallists and hobbyists, with posts ranging from a call for a general moratorium on AI to the view that AI plays too small a role in mathematics, while noting that contributors remain concentrated in Western Europe and North America, few have experience developing or evaluating LLMs, and only a handful of the forty-or-so published posts were written by women, and inviting
arXiv

ProgramDistill: Turning Interactive Web Apps into Verifiable Reference-Guided SWE Tasks

This work introduces the ProgramDistill benchmark and a fully automated mine–craft–patch pipeline that factorizes 26 interactive web applications into 1,975 replay-verified behaviors and constructs 4,063 tasks, letting coding agents infer and restore missing functionality by interacting with a reference application whose source is hidden; across nine frontier agents, the best full-application reconstruction success is 49.2%, and partial-application reconstruction success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.
arXiv

Routing Multiple Agents Below the Sum of Distances: A Parameterized Complexity Characterization of Transient Multiagent Pathfinding

This work studies Transient Multiagent Pathfinding, in which agents must be routed without collisions and disappear upon reaching their destinations, and shows that the problem is fixed-parameter tractable in the combined parameter k+ζ, where k is the number of agents and ζ=L−λ is the gap between the sequential-routing upper bound L=1+Σdist(si,ti) and the target makespan λ, running in 2^{O(k²ζ)}·n^{O(1)} time, complemented by matching lower bounds (W[1]-hardness for k alone, W[1]-hardness for ζ alone when terminals need not be distinct, and no polynomial kernel for k+ζ) and by fixed-parameter tractability in ζ alone when all terminals are pairwise distinct, running in 2^{O(ζ³)}·n^{O(1)} time.
arXiv

AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time

This work introduces AdaGeoVLN, a streaming vision-language navigation framework that couples VGGT geometry-foundation-model representations at depths 11, 17, and 23 to the first three Qwen3.5-4B decoder layers and retains historical VGGT global-attention KV states under a fixed budget according to instruction relevance, geometric confidence, and transition novelty, achieving 55.7%/51.4% and 54.1%/44.7% SR/SPL on R2R-CE and RxR-CE Val-Unseen with a single RGB stream and no additional navigation-specific external data, with ablations showing that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations and that bounded navigation-aware retention preserves navigation performance while reducing GFM-KV memory.
arXiv

CERA-MoA: A Mixture-of-Agents Framework Where Routing Mechanisms and Continually Learning Agents Co-Evolve

This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.
NVIDIA Research

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.
arXiv

TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

The work presents TeleAntiFraud 2.0, a refreshable Chinese call-audio benchmark organized as monthly frozen snapshots, built with a Mixed-Tree Anti-Fraud Generation Pipeline that turns online fraud-case abstracts into profile-grounded scenarios and expands them into fraud and near-domain lawful sibling dialogues sharing context and diverging only at label-bearing actions, rendered as role-matched speech; each frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), controlled text experiments show three classifiers reach perfect Macro-F1 against unrelated or ordinary negatives but drop to 0.65-0.68 with near-domain sibling negatives, and full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity.
arXiv

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

The work presents a multi-step framework that first corrects word segmentation with ChatGPT-4o, then fills missing entities via context-free lexicons with a minimum frequency of three, and finally applies frequency-based iterative self-training with a dual threshold on logits probabilities and the 90th percentile of self-attention scores, using a multilingual XLM-RoBERTa-large model to select candidates by F1; on manually revised validation/test splits for Urdu MK-PUCIT, Shahmukhi (Western Punjabi), and Sindhi SiNER, fine-tuning XLM-RoBERTa-large yields test micro-F1 gains of 3.96, 1.40, and 1.44 points, while ChatGPT-4o zero/few-shot NER remains below the supervised model.
arXiv

Clueing Up LLMs with Tool-Augmented Deductive Reasoning

This work adapts the board game Clue into a text-based multi-agent environment with six LLM players (three each from GPT-4o-mini and Gemini-2.5-Flash) across 18 games and introduces an external possibility-matrix tool (YES/NO/MAYBE cells plus an accusation_ready signal) that externalizes belief-state tracking, yielding near-perfect accusation accuracy for tool-augmented agents (Gemini-2.5-Flash 1.00, GPT-4o-mini 0.96), raising non-tool agents in mixed games from about 0.81/3 to about 2.71/3, while leaving agents' autonomy over accusation timing unchanged.
arXiv

Which LLM is Best for Translating Natural Language Goals to PDDL

This paper designs an iteratively refined prompt template that lets six contemporary large language models translate informal video-game testing goals written in natural language into PDDL goals for classical planning, and systematically evaluates them on 90 natural language goals (45 expressible and 45 inexpressible across eight domains) for correctness, speed, and error tendencies, finding that all models exceed 92% correctness, with Gemini 2.5 Flash highest at 96% and fewest false positives, while GPT-4.1 is fastest.
NVIDIA Research

Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers

Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a coalition convening the full AI and power value chain around a technology-neutral, performance-based approach that lets data centers dynamically manage electricity use to speed interconnection, strengthen reliability and protect affordability.
OpenAI

Reimagining Advertising with AI: OpenAI Introduces Sponsored Agents and Marketing Tools

OpenAI published an article introducing its AI-powered advertising experiences, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify, aiming to explore new forms of advertising in the AI era.
MIT Technology Review

Building the materials foundation for AI: Syensqo on advanced materials and AI as a two-way driver

This MIT Technology Review Insights conversation produced in partnership with Syensqo records the views of Mike Finelli, Syensqo's chief technology and innovation officer and chief North America officer: AI is pushing semiconductors and data centers toward physical limits, which piles up more simultaneous requirements on advanced materials, and Syensqo is responding by developing materials for high-voltage data center architectures, advanced sealing materials for semiconductor manufacturing, and thermal-management solutions including direct immersion cooling fluids, while working with Microsoft on AI agents that digitally synthesize millions of candidate molecules, predict their performance through physics-based simulation, and rank them down to roughly a hundred candidates for laboratory
MIT Technology Review

The Download: AI's Trillion-Dollar Gamble and OpenAI's Biology Data Bid

This is an edition of MIT Technology Review's daily newsletter The Download, rounding up technology news, with highlights including a financial analysis that starts from hyperscalers' nearly $1.1 trillion in data-center spending through 2027 and asks how fast their earnings must grow to break even by 2030, and the OpenAI Foundation's announcement that it will fund policy analyst Ruxandra Teslo's idea of obtaining regulatory filings and safety data from failed biotech companies through bankruptcy proceedings to build what she calls "biotech's lost archive."
OpenAI

How to connect AI usage to business value

This OpenAI article explains how ChatGPT Work and Codex analytics help teams understand AI usage and spend, identify training needs, and connect adoption to business outcomes.
Mistral AI

Mistral and Mozilla Are Bringing Open, Private and Multilingual AI to Your Web Browser

Mistral and Mozilla announced a partnership in which Mistral models now power Firefox's AI browsing assistant Smart Window (beta), starting with users in France and North America and expected to reach the United Kingdom and Germany later this year, while emphasizing open distribution, fine-tuning on regional languages and dialects, conversations not saved on Mozilla's servers by default with zero data retention, and bringing sovereign AI to everyday users.
OpenAI

OpenAI Economic Research: How Workers Are Unlocking New Ways of Working

An analysis from New OpenAI Economic Research indicates that workers are using AI for purposes beyond their traditional job roles, and identifies which of these new activities become recurring parts of their work.
NVIDIA Research

University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

Manchester physicists David Topping and Hao Zhang, working with the NVIDIA Earth-2 team, moved generative models originally built for weather into air quality forecasting: they trained the Earth-2 CorrDiff generative downscaling model on a year of hourly UK chemistry-climate simulation data, completing training in two days on a single eight-GPU node of Isambard-AI to produce a UK-wide pollution model at 2-3 square kilometer resolution, added Earth-2 StormCast for time-dependent forecasts that directly use air quality observations, demonstrated inference and smaller training runs on the DGX Spark desktop AI system, and plan to release open source training data and workflows so other countries and regions can train their own models.
ORBi UMONS

GHAgentFiles: A Dataset of Coding Agent File Histories in GitHub Repositories

This work builds and releases the GHAgentFiles dataset together with the open-source extraction tool cofee, identifying 27,717 repositories with coding agent files out of 165,281 candidates and extracting 142,294 context, skill and subagent file histories (401,873 file revisions across 185,795 commits, spanning 22 November 2022 to 1 July 2026), and it presents a preliminary observational analysis showing that coding agent files emerge around early 2025, peak in additions and modifications in early 2026, and that the generic AGENTS.md has become the most popular naming practice.
Diagnosis (Berlin, Germany)

Early evidence for a multi-agent AI simulator for clinical reasoning practice: performance, consistency, and challenges

In Fall 2024, 175 second-year medical students completed three MAESSCR multi-agent LLM clinical encounters as coursework, and six clinician-educators rated 120 randomly sampled transcripts with a dichotomous tool, finding 92% (110/120) adherence to scripted details, 2.5% (3/120) diagnosis-changing information, 12% (14/120) unrealistic patient portrayal, and 17% (20/120) technical issues, while the most prominent problem was agents interpreting findings before students could, occurring in 28% (34/120) of encounters for history/physical exam agents and 59% (71/120) for diagnostics/management agents, with most disruptions judged minor.
Journal of medical imaging and radiation sciences

Radiomics and its application to neuro-oncology: A narrative review of advances, clinical application and implementation challenges

This narrative review searched PubMed, Scopus and Web of Science (2012-2024) to map the role of radiomics and artificial intelligence in neuro-oncology across diagnostic, prognostic and therapeutic potential, reporting that radiomics can automatically extract quantitative features from MRI, CT and PET/CT that are not visible to the human eye, has been shown useful for pre-surgical classification of gliomas, survival prediction and differentiation between tumour progression and pseudo-progression, may improve diagnostic accuracy and support patient stratification by molecular biomarkers such as IDH and MGMT when integrated with deep learning, and links image phenotypes to genetic alterations through radiogenomics to enhance personalised medicine, while limitations persist from lack of proto
Science advances

Microwave diffractive neural network chips: sensing and computing on a millimeter-scale GaAs chip

This work fabricates a chip-scale microwave diffractive neural network (MDNN) in a GaAs semiconductor process, integrating cascaded couplers and phase shifters to implement a diffraction network within a millimeter-scale footprint, reducing the size of conventional MDNNs by over four orders of magnitude, achieving a computational latency of 2.05 ns and a system-level energy efficiency of 0.83 TOPS/W, and reaching more than 86% accuracy across three functional prototypes—MNIST handwritten digit recognition, multi-user interference suppression, and real-time obstacle perception for drones—thereby validating the chip's capability to directly perform both digital image processing and in-situ electromagnetic information processing in the microwave domain.
JMIR Medical Education

The AWARE Framework: An Educational Interview Framework for Assessing Patients' Use of Conversational AI in Mental Health Care

This article introduces the AWARE (AI use, why, attachment, reality and risk, and effect on functioning) framework as an educational tool to help mental health professionals systematically assess patients' use of conversational AI through five clinically relevant domains, and discusses incorporating it into undergraduate, postgraduate, and continuing professional education along with priorities for future research.
medRxiv

An AI Competency Framework for Emergency Medicine: A Multiphase Consensus Process

Using a Nominal Group Technique session (May 2025, n=6) and a two-round modified Delphi process (April–May 2026, expert panel n=13 across 12 academic medical centers, with 77% Round 1 and 100% Round 2 response rates), this study developed the first specialty-specific AI competency framework for a United States medical specialty, comprising 5 themes (Communicating about AI, Understanding appropriate use cases, Interacting with AI, AI risk management, Cognitive impacts of AI), 5 derived competencies (one-to-one theme-to-competency mapping endorsed by 12 of 13 panelists), 19 subthemes, and 10 retained clinical scenarios, with all themes, derived competencies, and individually rated subthemes meeting pre-specified consensus thresholds and 9 of 10 scenarios reaching consensus while 1 was retain
Clinical therapeutics

Medical Device Safety Throughout the Product Lifecycle: Regulatory Frameworks, Digital Transformation, and the Role of Artificial Intelligence in Devices and Monitoring

This narrative review, based on structured searches of PubMed and ScienceDirect (2016-2026, plus a clinical-research-focused search for 2013-2026) supplemented by FDA, EMA, Cochrane Library, and industry sources, examines medical device safety across the full lifecycle from early design and clinical investigation through regulatory review, postmarket surveillance, software change management, cybersecurity, and AI/ML-enabled device oversight, arguing that device safety requires a broader socio-technical framework than drug safety and that safety evidence is shifting from retrospective, passive reporting toward proactive, real-time, lifecycle-based generation.
JMIR Medical Informatics

Prediction Models for In-Hospital Delirium Using Routinely Collected Electronic Health Record Data: Systematic Review

This systematic review searched PubMed, MEDLINE, Embase, PsycINFO, and Web of Science from inception to November 11, 2025, and included 29 studies that developed, validated, or evaluated multivariable prediction models using routinely collected electronic health record or administrative data to predict acute mental status deterioration during adult hospital admissions, all operationalized as delirium; using CHARMS and TRIPOD/TRIPOD-AI for data extraction and PROBAST for risk of bias, it found that the evidence clustered into four overlapping prediction tasks (admission or early-stay risk stratification, perioperative or postoperative prediction, dynamic intensive care unit prediction, and external validation or workflow evaluation of existing tools), that most studies were retrospective co
ACM Computing Surveys

Mixed-Precision Quantization for Language Models: Techniques and Prospects

This survey organizes the landscape of mixed-precision quantization frameworks for language models (MXPLMs): it first reviews quantization fundamentals including uniform and non-uniform quantizers, quantization granularity, and widely used post-training quantization methods, then categorizes and compares recent MXPLM frameworks by their bit allocation strategies and precision configurations across weights, activations, and key-value caches, contrasts them with earlier mixed-precision methods for deep neural networks to identify strategies that transfer and those that face challenges in the LM setting, and closes with open issues such as hardware-aware design, activation quantization, and scalable optimization for billion-parameter models.
发表出处待核验

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

This work introduces information satisfaction—a query paired with a reader persona—as an axis of summarization evaluation, and through five perturbation tests plus an expert human evaluation finds that traditional metrics such as ROUGE and BERTScore and LLM-as-judge metrics such as Llama-3.3-70B and Prometheus-7B mostly fail basic perturbation checks and agree with reader preferences at near-chance levels, indicating that existing metrics are insufficient measures of how well a summary serves a specific reader's informational needs.
Figshare

Replication package — The Institutional Window: How Contract Law Bounds Liability Signaling of Human Fallback Capability under Generative AI (v1.3.1.1)

This replication package supplies full verifiable materials for Bauer (2026), which asks when a liability commitment can still certify a provider's preserved human fallback capability once generative AI makes the output itself uninformative, mapping a posted cap and agreed-damages term into retained exposure through four legal primitives and deriving the message set {0} u [F, C] in which low types pool at zero, intermediate types separate on a schedule anchored at F, and high types may pool at the ceiling, with separation beginning at the bottom type where law removes the zero-exposure region.
Biologia futura

Microbial diversity: the essential foundation for life on our planet

This review states that microbial diversity underpins human health, agricultural productivity, ecological balance, and ecosystem functioning, and it surveys the role of the gut microbial community in immune regulation, metabolism, and disease prevention, the contributions of interactions among plants, fungi, bacteria, and other soil microorganisms to carbon sequestration, nutrient cycling, stress resilience, and sustainable agricultural productivity in terrestrial ecosystems, and emerging microbiome-based therapies such as precision probiotics, postbiotics, faecal microbiota transplantation, and personalized microbiome medicine, while noting that multi-omic techniques, synthetic microbial genomes, microbiome engineering, and artificial intelligence enable emerging uses in agriculture, envi
Climatic Change

Generative Debunking of Climate Misinformation: Automatically Producing 'Truth Sandwich' Rebuttals with Large Language Models

This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.
Current oncology reports

Recent Advances in Surveillance Strategies for Nasopharyngeal Carcinoma: From Guideline Follow-up to Individualized Precision Surveillance

This review systematically examines current follow-up protocols and recent developments for nasopharyngeal carcinoma (NPC): on one hand it compares and analyzes major follow-up guidelines in terms of follow-up frequency, imaging modalities (MRI and PET), plasma EBV-DNA monitoring, and functional assessments; on the other hand it elaborates on the application prospects and research progress of genomics, radiomics, and artificial intelligence in NPC surveillance, noting that risk-stratified, individualized follow-up strategies such as those based on conditional survival models can enhance the cost-effectiveness of surveillance, while radiomics and artificial intelligence show promise for improving recurrence risk assessment, prognostic stratification, and individualized surveillance.
BMC Geriatrics

Quality and Safety of Large Language Model–Generated Medication Review Outputs in Geriatric Pharmacotherapy: A Two-Stage Comparative Vignette-Based Benchmark Evaluation

Using 20 standardised geriatric pharmacotherapy vignettes (fictional older adults aged 72–88 across four clinical domains, each containing three potentially inappropriate medications and one START-type omission anchored to the AGS Beers Criteria and STOPP/START version 3), this study had GPT-5.2, Claude Sonnet 4.5, and Gemini 3 Pro respond under an identical master prompt and default end-user settings, with two geriatricians blinded to model identity independently rating anonymised outputs on a 100-point rubric (output quality 0–80, critical safety-risk prioritisation 0–20), finding that Stage 1 item-level answer-key concordance was uniformly high with limited between-model discrimination while expert-rated total scores differed significantly across models (p < 0.001; Kendall's W = 0.
Chemical Society reviews

Cathode lithium-rich compensators for next-generation lithium-ion batteries: classification framework, challenges, and synergistic prelithiation strategies

This review addresses active lithium loss during the initial cycling and long-term operation of lithium-ion batteries by establishing a classification framework for cathode lithium-rich compensators (LRCs) based on lithium compensation mechanisms, covering binary, ternary, over-lithiated, sacrificial lithium salt, and sustained-release types; it systematically summarizes their working principles, typical charge compensation pathways, and practical performance, discusses key challenges including air instability, gas evolution, high delithiation voltage, and processing issues, highlights mitigation strategies involving nanoscale engineering, surface coating, defect and doping design, and electrolyte optimization, and further proposes a synergistic multimodal lithium-compensation strategy int
JMIR Formative Research

Problematic Reliance on Generative AI in an Anxious Young Adult: A Case Report

This case report describes a woman in her mid-20s with generalized anxiety disorder and major depressive disorder and a history of strong social, academic, and occupational functioning who developed a pattern of functional dependence on ChatGPT, outsourcing routine cognitive and interpersonal tasks such as composing emails, interpreting social interactions, predicting the future, and making decisions, and becoming increasingly uncomfortable completing such tasks independently; the authors frame this as cognitive offloading, reduced confidence in independent judgment, and reinforcement of externalized thinking using the I-PACE model, and suggest that the unlimited accessibility of AI tools may intensify reassurance seeking and worsen tolerance of uncertainty.
Mendeley Data

A corrosion inhibitor dataset built from 5,597 publications

This work assembles a corrosion inhibitor dataset from 5,597 publications using a multi-agent extraction pipeline, releasing a manually verified IE_datasets.xlsx, LLM-extracted json_extracted.zip, and IE_PH_4_materials.xlsx for model construction, with fields covering reference metadata, inhibitor name and composition, anodic/cathodic/mixed type, film mechanism, SMILES, corrosion material name, grade, composition, processing and heat treatment, corrosive medium, medium type, concentration and temperature, and test temperature, test time, inhibitor concentration, test method and inhibition efficiency percentage.
ACM Journal on Computing and Sustainable Societies

The Unbearable Lightness of Prompting: A Critical Reflection on the Environmental Impact of genAI use in Design Education

Using a 2023 workshop with 49 students as a motivating example, this paper critically reflects on the energy costs of using genAI in design education and develops a set of five alternative stances, with related actions, to support the conscious use of genAI in design education.
Bioscience Reports

Epigenetic Mechanisms and Clinical Translation in Ovarian Cancer: From Molecular Pathways to Precision Therapy

This review systematically examines how DNA methylation, histone modifications, chromatin remodeling, and non-coding RNA networks drive chemoresistance and recurrence in ovarian cancer, summarizes clinical trial results of epigenetic agents (DNMT, HDAC, and EZH2 inhibitors) combined with PARP inhibitors or immunotherapy, and reviews biomarker advances based on circulating cfDNA methylation, circulating miRNAs, and AI-based liquid biopsy platforms, proposing a precision oncology framework for patient stratification and real-time monitoring of chemoresistance.
Neurotherapeutics : the journal of the American Society for Experimental NeuroTherapeutics

Accelerating discovery: Transformative clinical trial models in neuro-oncology

This review proposes and organizes a framework of emerging clinical trial models for central nervous system tumors, including master protocol designs, Bayesian adaptive frameworks, trials as active discovery platforms embedding longitudinal tissue sampling, window-of-opportunity designs and multi-omic profiling, and decentralized models with artificial intelligence tools, arguing that trials should be reimagined as dynamic, biologically integrated, learning-based systems rather than static tests of individual agents in order to accelerate therapeutic progress in neuro-oncology.
Trends in pharmacological sciences

Closing AI drug-regulatory gaps through harmonized oversight

The article notes that while artificial intelligence accelerates drug discovery it also creates regulatory gaps, with more than 100 AI-assisted pipelines in trials and frameworks lacking enforceable standards; it therefore proposes a risk-tiered framework that distinguishes discovery AI from evidence-generating AI, mandates impact assessments and Investigational New Drug disclosure when AI influences decisions, and transforms guidelines into binding, risk-proportionate regulation.
European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology -

Artificial Intelligence in Otolaryngology: Current Applications, Limitations, and Future Perspectives

This narrative review, based on a structured search of PubMed/MEDLINE, Scopus, and Web of Science, maps the clinical applications of artificial intelligence across otolaryngology subspecialties, noting that deep learning shows potential in sinonasal disease detection, automated image segmentation, lymph node metastasis prediction, thyroid nodule classification, and prognostic modeling in head and neck cancer, while multimodal systems and generative large language models are emerging in medical education, image interpretation, differential diagnosis, and clinical decision support; however, limited external validation, retrospective designs, dataset heterogeneity, algorithmic bias, lack of transparency, privacy concerns, medico-legal uncertainty, and automation bias still constrain broad imp
bioRxiv

The Unreasonable Effectiveness of Cell Types in Describing Neuronal Physiological Features

Using paired transcriptomic and electrophysiological patch-sequencing data from 495 human neurons from neurosurgical tissue, this study compared how well traditional transcriptomic cell type classification, representations from the foundation model scGPT pretrained on large-scale scRNA-seq datasets, ion channel-coding genes, and highly variable genes predict electrophysiological features, finding that cluster-level cell type representations generally outperform highly variable gene selection, ion channel gene selection, and context-enriched scGPT embeddings, with the best results obtained by combining the outputs of separate cell type and scGPT-based models.
medRxiv

Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes

Using a zero-shot large language model pipeline (Gemini 3.5 Flash and Claude Haiku) to extract 46 symptoms from 2,728 discharge notes of 1,507 colorectal cancer patients in MIMIC-IV, this study built patient-level symptom co-occurrence networks with phi correlation (≥0.10) and Louvain community detection; both models converged on three clinically coherent symptom clusters — Systemic, CRC Disease-Specific, and Gastrointestinal — and Systemic cluster burden was associated with in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41) while CRC Disease-Specific cluster burden independently predicted 30-day readmission (OR=1.20), with both associations robust to adjustment for metastatic disease.
medRxiv

Task-Specific Quality Gating for Retinal OCT B-Scans: Learned Representations Over Scalar Metrics in Choroid Segmentation

Using choroid segmentation as a prototype task on 6,076 OCT B-scans from 80 subjects, this study systematically compared scalar no-reference image quality metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), general-purpose ImageNet-pretrained representations, and retinal foundation models as task-specific quality gates, finding that scalar metrics correlate weakly with segmentation Dice (|r| < 0.20), that general-purpose pretrained representations reach linear-probe ROC-AUC up to about 0.77, that the OCT-specific foundation model RETFound reaches about 0.81, and that only the retinal foundation model embeddings form quality-aligned unsupervised K-Means clusters exceeding a patient-level permutation null.
RNA Biology

HyLnc: A Hybrid Framework Combining Deep Contextual Embeddings and Handcrafted Sequence Features for Long Non-Coding RNA Prediction

The study proposes HyLnc, a framework that first pre-trains a custom BERT model on a large corpus of metazoan RNA sequences with masked language modelling, then fine-tunes it on curated lncRNA and protein-coding transcript datasets to extract 256-dimensional deep embeddings, while computing 348 handcrafted features including ORF characteristics, UTR properties, nucleotide composition and Fickett scores; after multi-stage feature selection, multiple machine learning classifiers were evaluated, with random forest performing best and achieving 91.30% accuracy, 91.23% F1-score and 82.60 MCC on an independent validation dataset, outperforming several existing lncRNA prediction tools.
Aesthetic Surgery Journal

Current Concepts and Emerging Technologies in Aesthetic Outcome Assessment of Breast Reconstruction: A Systematic Review

This systematic review searched studies from 2000 to 2025 evaluating aesthetic outcomes after implant-based, autologous, or hybrid breast reconstruction, included 51 studies with 7711 participants from 16 countries, classified assessments as subjective or objective, found subjective tools led by BREAST-Q (35/51, 69%) and objective methods in 19 studies including 3-dimensional surface imaging (8/51, 16%), BCCT.core (7/51, 14%), eye tracking (3/51, 6%), and artificial intelligence-based analyses (3/51, 6%), and noted that no single modality comprehensively addresses all aesthetic domains, with a progressive shift toward multimodal evaluation.
bioRxiv

nanorepertoire: an end-to-end Nextflow pipeline for nanobody repertoire analysis

This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.
Heart & Lung

Multimodal Large Language Models in Prehospital ECG Triage for Emergent Catheterization Laboratory Activation: A Retrospective Comparative Analysis

This study retrospectively analyzed 615 ECGs from 270 emergency medical service patient encounters with concern for acute myocardial infarction, using cardiology activation of the STEMI pathway as the reference standard, and compared three multimodal large language models with an ECG machine algorithm, finding that Gemini had the highest sensitivity (95.3%) but extremely poor specificity (9.4%), ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%), while the ECG machine algorithm was more balanced with sensitivity of 67.7% and specificity of 64.2%, suggesting that general-purpose large language models are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive emergency workflows.
bioRxiv

Corpus-wide causality: Algorithm design & application for aggregating gene-disease causal evidence

This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
medRxiv

Time-resolved predictability of end-of-therapy outcome and relapse after cure in Phase 3 tuberculosis trials

Using harmonised clinical data from two Phase 3 trials (2,918 participants), this study trained monthly tabular models from baseline to therapy end for time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse, finding that EOT outcome prediction improved after month 3 (ROC-AUC up to 0.84, driven by sputum-smear and solid culture), whereas relapse prediction among those with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most; large language model-derived embedding models matched tabular relapse models throughout therapy and outperformed tabular EOT models at months 3 and 4 (ΔROC-AUC 0.12 and 0.
Pharmacology & therapeutics

Drug Resistance in Cancer: A Conceptual Shift from Single-Gene Mechanisms to Systemic Adaptive Resistance

This review proposes reframing multidrug resistance (MDR) from conventional single-gene mechanism models toward systemic adaptive resistance as a complex and evolving biological phenomenon, and systematically reviews classical resistance mechanisms (drug uptake/efflux, compensatory pathway activation, apoptosis evasion), non-genetic resistance mechanisms (drug-tolerant persisters, DTPs; epithelial-mesenchymal transition, EMT; cancer stem cell, CSC, plasticity and related factors), how interactions among these mechanisms contribute to tumor persistence and MDR evolution, and translational barriers together with emerging pharmacological strategies including longitudinal molecular monitoring and artificial intelligence (AI)-assisted drug resistance prediction.
medRxiv

Do Large Language Models Use the Clinical Vignette? A Question Ablation Study on the Orthopaedic In-Training Examination

This study ablated question components across 792 Orthopaedic In-Training Examination (OITE) questions from 2020 through 2024 (434 with clinical images, 358 without), evaluating three open-source Ministral-3 models (3B, 8B, 14B) and five proprietary models (Claude Haiku-4.5, Sonnet-4.6, Opus-4.8, GPT-5.6 Luna, GPT-5.6 Terra), and found that pooled accuracy on image-containing questions was 59.13% (58.21-60.08) with complete information and 59.13% (58.18-60.11) without images, dropping to 49.05% (48.07-50.00) without the clinical vignette; non-image questions fell from 72.94% (72.10-73.85) with complete clinical context to 53.53% (52.41-54.68) without the vignette and 37.36% (36.28-38.48) with answer options alone; with options only, every model exceeded the 25% random baseline (highest 45.
arXiv

Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

The work proposes dynamic context adaptation: a validation-generation loop in which a validation agent extracts structured diagnostic feedback from execution traces, a generation agent proposes multiple candidates per iteration, a knowledge graph supplies semantic constraints, and simulated annealing performs non-greedy selection, targeting the LLM code-generation limitation the authors call static binding; across eight problems the method outperforms zero-shot, Reflexion, and OpenEvolve on seven of eight problems at both 300 and 600 evaluations (p < 0.01), achieves the best score at 1000 evaluations on the primary motivating problem of cross-coupled optimization (0.694 vs. 0.681, p = 0.019, d = 0.52), and ablations identify structured execution feedback as the primary driver.
arXiv

GenFT: A Generative Parameter-Efficient Fine-Tuning Method for Pretrained Foundation Models

This work proposes GenFT, a W0-conditioned parameter-efficient fine-tuning method in which a deterministic generator produces task-specific updates ΔW by applying row and column transformations to the pretrained weights W0 together with a shared-specific decomposition, reporting competitive or better average performance on GLUE (RoBERTaBase, 85.87% average with 0.24M parameters), VTAB-1K (ViT-B/16, 74.50% average with 0.27M parameters) and FGVC (90.38% average with 0.29M parameters), plus a perplexity pilot study on LLaMA-7B with an Alpaca subset.
Journal of the European Meteorological Society.

Machine learning is revolutionizing weather forecasting — the next step is a change in how we work

This article by Peter Dueben, Peter Bauer, Oliver Fuhrer, Nikolay Koldunov and Jørn Kristiansen argues that, following the success of machine learning in producing weather predictions with competitive skill compared to complex traditional systems, attention should shift from forecast output to the working practices that make prediction systems possible, and that machine learning and recent digital technologies will reshape the forecasting value chain — how models are coded and developed, how observations and Earth-system data are exploited, how data and computing are managed, how systems are verified, and how information is created, evaluated and turned into services; it discusses six non-exhaustive areas in which agentic software engineering, open and compressed data, shared verification
Journal of Medical Internet Research

Dynamic Prediction of 30-Day Mortality in Patients With Trauma Using a Hybrid Neural Network Model: Model Development and Evaluation Study

Using electronic health record data from 9,496 patients with trauma treated in the Capital Region of Denmark between 2017 and 2024, this study developed a hybrid neural network combining tabular and sequential data to predict 30-day all-cause mortality at any time point from prehospital care to discharge, achieving AUROC 0.962 and AUPRC 0.655 on a holdout set of 1,829 patients and AUROC 0.905 at 1 hour from first patient contact in active-cohort evaluation, with better discrimination than the Revised Trauma Score (mean ΔAUROC +0.297) and the Trauma and Injury Severity Score (mean ΔAUROC +0.169).
Transfusion clinique et biologique : journal de la Societe francaise de transfusion sanguine

Infectious risks of transfusion: a 10-year look back and perspectives

This review reflects on how infectious diseases shaped transfusion services from 2016 to 2026, covering innovations in donor selection, testing and pathogen reduction, notable pathogens including Zika virus, Plasmodium, Babesia and SARS-CoV-2, and favorable developments such as individualized risk assessment, relaxation of donor deferral policies, and AI and machine learning tools for surveillance and horizon scanning, while noting that many of these gains do not extend to low and low-middle income countries.
International Journal of Cardiology

AI-assisted handheld echocardiography for hospital bedside cardiac triage: the prospective OPTIMUST implementation study

OPTIMUST was a prospective, single-center implementation study in which physicians without formal echocardiography certification completed a structured two-month curriculum and performed focused examinations using Caption AI-enabled handheld ultrasound on cardiology and non-cardiology wards; of 287 attempted examinations, 206 (71.8%) were analyzable, operator assessment correlated with expert review of the same handheld image sets for LVEF (r = 0.84), filling-pressure category agreement gave a quadratic weighted κ of 0.659, physicians reported that handheld findings changed or confirmed management in 93.7% of examinations, and 32.8% underwent comprehensive echocardiography within one month.
medRxiv

Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports Using Large Language Models

Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Clinical transplantation and research

Artificial intelligence in preclinical nonhuman primate xenotransplantation: bridging the gap from data complexity to clinical precision

This review synthesizes current evidence on the use of artificial intelligence and machine learning in preclinical nonhuman primate xenotransplantation, noting that computer vision can support continuous noninvasive behavioral phenotyping and pain assessment, anomaly detection algorithms can extract early warning signals from biosignal streams, multiomics integration can identify xenograft-specific biomarker signatures, digital twin frameworks may enable hypothesis-generating simulations of prospective human recipient responses, and explainable AI can make model outputs more transparent and regulatory defensible, and it proposes a research agenda through which AI-augmented nonhuman primate experimentation can become a bridge from preclinical data complexity to clinical precision in xenotra
发表出处待核验

When AI Says "I have been in similar situations": Synthetic Lived Experience in Peer-like Caregiver Support

In the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD), this work compares caregiver support exchanges from online communities with peer-like responses prompted from three LLMs (LLaMA, GPT-4o-mini, and MedGemma), using psycholinguistic and qualitative analysis to show that peer responses used significantly more first-person and past-focused language than peer-like AI responses, identifies seven types of personal narratives in human peer support, and finds that AI often captures their emotional work while potentially fabricating experiential grounding, thereby naming a narrative authenticity gap and a synthetic lived experience paradox.
arXiv

Geometric Self-supervised Pretraining on 3D Protein Structures Using Subgraphs

This work proposes a new self-supervised pretraining task for 3D graph neural networks: pretrained on 542k SwissProt protein structures from the AlphaFold Database, the model predicts the Euclidean distance between the geometric centroid of each protein subgraph (2-hop ego networks centered on 10% of amino acids) and the global geometric centroid of the whole protein, discretized into 10 equal bins and trained with cross-entropy; across ProNet, SchNet, and GCN backbones and ca_base, ca_angles, and ca_bb featurizations, it improves Fold, Superfamily, Family, and React classification by up to about 6% over no-pretraining and edge-distance pretraining baselines, without multiple views, augmentations, or masking strategies.
The Journal of Clinical Endocrinology & Metabolism

Artificial Intelligence-Enabled Analysis of Radiology Reports: Epidemiology and Outcomes of Incidental Thyroid Findings

This study developed and deployed a transformer-based natural language processing pipeline to identify incidental thyroid findings in radiology reports from 115,683 adults without prior thyroid disease across Mayo Clinic sites from July 1, 2017, to September 30, 2023, finding that 7.8% had such findings (92.9% nodular) and that these findings were associated with higher odds of downstream thyroid nodule diagnosis, biopsy, thyroidectomy, and thyroid cancer diagnosis, with most cancers being papillary.
Trends in biotechnology

Liquid biopsy in glioblastoma: emerging technologies and translational opportunities

This review focuses on liquid biopsy for glioblastoma, noting that tissue biopsy is invasive and fails to capture tumor heterogeneity or temporal dynamics, while liquid biopsy can provide noninvasive real-time monitoring via circulating biomarkers such as cell-free DNA, circulating tumor DNA, circulating tumor cells, and extracellular vesicles in plasma, cerebrospinal fluid, urine, and saliva; it advances beyond prior biomarker catalogs by delivering a quantitative technology scorecard comparing cfDNA-, CTC-, and EV-based platforms in terms of sensitivity, clinical actionability, cost, and scalability, and introduces a multimodal decision matrix and an AI-driven fusion pipeline integrating fragmentomics, EV proteomics, and CTC transcriptomics to enhance minimal residual disease detection a
medRxiv

Preoperative Social Connection and Postoperative Outcomes in Adults Undergoing Surgery: A Systematic Review and Meta-analysis

This systematic review and meta-analysis of 445 studies found that weaker preoperative social connection was associated with early postoperative mortality (OR 1.50, 95% CI 1.12-2.01) and non-home discharge (OR 1.95, 95% CI 1.35-2.81), while confidence intervals for postoperative survival, unplanned readmission, and complications all included 1, with overall certainty ranging from low to very low.
arXiv

Re-evaluating the Advancements of Heterophilic Graph Learning: A Three-Way Dataset Taxonomy and Quantitative Evaluation of Homophily Metrics

This work fine-tunes baseline models on 27 widely used benchmark datasets, categorizes heterophilic datasets into malignant, benign, and ambiguous groups based on whether graph-aware models underperform their coupled graph-agnostic counterparts, re-evaluates 11 SOTA heterophily-specific models, and conducts the first quantitative evaluation of 11 homophily metrics on synthetic graphs from three generation methods, finding that most SOTA models do not significantly outperform the best baselines and that classic metrics remain competitive.
Biochimie

MAPK signalling landscape of Leishmania: structural diversity, biological functions and therapeutic potential

This review synthesizes the 15 MAPK homologues across pathogenic Leishmania species, linking individual kinases to parasite differentiation, intracellular survival, stress tolerance, motility, virulence and drug response, distinguishing experimentally validated mechanistic targets from computationally proposed candidates, and evaluating small-molecule and natural-product inhibitors, drug repurposing, structure-guided approaches, and emerging contributions from molecular dynamics, AlphaFold-based modelling, artificial intelligence and nanotechnology, while noting evidence for MAPK10 as a vaccine-associated antigen.
JMIR AI

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model-Generated Medical Information

This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
Advanced science (Weinheim, Baden-Wurttemberg, Germany)

AI-Guided Phenotypic Drug Repurposing Against Streptococcus pneumoniae

This study applied AI-guided phenotypic drug repurposing to drug-resistant Streptococcus pneumoniae, using ensembles of transformer, graph, and tree models trained on 1849 actives and 34 503 inactives to prospectively examine 6747 drugs, selecting 11 candidate antibiotics of which nine strongly reduced in vitro growth of S. pneumoniae R6 (IC50 ≤ 0.4 µg/mL), with the most potent drugs thiostrepton and ceftiofur showing IC50 values of 0.0001 µg/mL (60.1 pM) and 0.0004 µg/mL (764 pM), respectively, and thiostrepton remaining highly potent against multidrug-resistant strains.
Structure (London, England : 1993)

AlphaBridge: Tools for the analysis of predicted biomolecular complexes

This work presents AlphaBridge, a reproducible, objective, and automated toolkit, also available as a web server, that combines AlphaFold3 confidence metrics to cluster sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes, visualizes interaction interfaces within confidence limits via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions linked to interactive graphics, and was validated for scoring binary and multi-component protein complexes with real-life examples discussed.
The Permanente journal

Can Artificial Intelligence Deliver in Real-World Health Systems? Early Insights From AIM-HI's 5 Funded Projects

This article synthesizes early cross-project insights from the 5 projects funded by the Augmented Intelligence in Medicine and Healthcare Initiative (AIM-HI), led by Kaiser Permanente and funded by the Gordon and Betty Moore Foundation and selected through a national, multistage review process using a structured scoring rubric, covering sepsis management, venous thromboembolism risk assessment, diabetic retinopathy screening, cardiac amyloidosis detection, and pediatric asthma risk prediction; it reports that real-world AI deployment was feasible across varied clinical environments, that common challenges included electronic health record integration, data complexity, regulatory requirements, and variation in clinical workflows, and that implementation success depends on thoughtful integra
medRxiv

Genomic foundation model-derived disruption profiling links somatic mutations to cancer biology and clinical outcomes

Using AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas, this study found that recurrent hotspot mutations showed substantially larger predicted protein-level effects while non-hotspot mutations exhibited larger regulatory effects across most cancer types; aggregating variant-level predictions into patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing yielded gene- and modality-specific profiles that reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden; among patients lacking recurrent hotspot mutations, higher predicte
Practical radiation oncology

Clinical Implementation and Evaluation of an Artificial Intelligence-Driven One-Click Automatic Planning System for Functional Lung Avoidance Radiotherapy

This study implemented AP-FLART, which integrates dosimetric score-based beam angle selection, multi-modality-guided dose prediction, and function-guided dose mimicking, within RayStation and evaluated it on a test dataset of 33 lung cancer patients who underwent SPECT ventilation or perfusion imaging and lung radiotherapy, finding that automatic FLART plans significantly reduced high-function lung mean dose by 15.1% versus manual conventional radiotherapy plans, lowered the probability of grade >=2 radiation pneumonitis by 6.25 percentage points (27%) among FLART-benefiting patients, achieved clinical benefits similar to manual FLART plans, were clinically acceptable without modification in 87.9% of cases, and cut planning time from 2-3 hours to approximately 8 minutes.
Journal of Nondestructive Evaluation

Slow Dynamics in Concrete: Effects of Temperature, Strength Variation, and Microcracking Damage

This study conditioned concrete prisms of four compressive strengths (f′c = 32–56 MPa) and two alkali-silica reaction (ASR) damaged specimens by compressive loading, monitored the subsequent velocity recovery with coda wave interferometry (CWI), and applied a self-referencing temperature correction to remove drift from ambient fluctuations of about ±0.2°C; for intact specimens the recovery rate mv and velocity drop magnitude |c| increased with strength (|c| from 3.08×10⁻⁴ to 6.02×10⁻⁴, mv from 6.77×10⁻⁵/s to 13.76×10⁻⁵/s) and recovery time shortened from 9.9 h to 6.
Research Square

LLMs for Survey Text Analysis: A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Using 903 open-ended responses across six variables from a European PhD student survey, this study had five human coders and GPT-5.4 each perform the same inductive content analysis procedure to produce codes and themes, and measured agreement with the Adjusted Rand Index (ARI), finding average human-LLM agreement of 0.61 for coding and 0.54 for themes, close to within-human consistency (0.68) and within-LLM consistency (0.76), with wide variation across variables and low within-entity consistency consistently accompanying low between-entity agreement.
RNA

A Continuum-Based Reaction-Diffusion Model Reveals Spatial Spread of Gene Silencing in Chromosomal Inactivation

This work develops a continuum-based reaction-diffusion model of XIST-mediated gene silencing spread on chromosomes, finding that XIST spread can be tuned by known negative feedback loops regulating its synthesis and degradation, while silencing spread is controlled by a wave-pinning mechanism driven by global regulation of the silencing complex together with local epigenetic regulators, and uses a 3D chromosome structure inferred from experimental data to show spatiotemporal regulation of silencing spread.
npj Digital Medicine

Auditing Sex/Gender Disparities in Emergency Triage with LLM-based Paired Comparisons

The study introduces a domain-agnostic paired-comparison approach that fine-tunes a large language model to emulate documented emergency triage decisions and then compares predictions on sex-swapped pairs in which only sex is flipped while documented clinical content is held constant, finding that otherwise identical presentations were more likely to receive a lower-severity (less urgent) predicted triage score as female than male, with a pooled per-pair rate of about 1.1% (95% CI 0.9–1.3) across more than 140,000 Bordeaux University Hospital admissions and a directionally consistent but larger 2.2% (1.7–2.7) in MIMIC-IV, while a model retrained on sex-neutralized inputs eliminated the between-sex prediction gap, indicating the asymmetry is mediated by explicit sex markers.
Anesthesiology

Development and External Validation of a Multimodal Artificial Intelligence Mortality Prediction Model of Critically Ill Patients Using Multicenter Data

Using 203,434 ICU admissions from 2001 to 2022 across more than 200 hospitals in the MIMIC-III, MIMIC-IV, eICU, and HiRID databases, the study developed and externally validated a multimodal deep-learning model that predicts subsequent inpatient mortality from time-invariant variables, time-variant variables, clinical notes, and chest x-ray images available within the first 24 h of ICU admission; with structured data alone the model reached an AUROC of 0.92 (95% CI, 0.90 to 0.93), an AUPRC of 0.53 (95% CI, 0.49 to 0.57), and a Brier score of 0.19 (95% CI, 0.18 to 0.20), external validation across eight eICU institutions yielded AUROCs of 0.84 to 0.92, and in the subgroup with both notes and imaging, adding text and images raised the AUROC modestly from 0.87 (95% CI, 0.85 to 0.89) to 0.
medRxiv

Tumour region identification guided scoring (TRIGS) and foundation model-based Tumour Infiltrating Lymphocyte scoring are prognostic for pathological complete response/event free survival in the triple negative patients in the PARTNER randomized controlled trial

This study updated an automated tumour infiltrating lymphocyte (TIL) assessment pipeline, proposing TRIGS aligned with clinical scoring guidelines and foundation model-based SAM-TIL, and showed in 166 neoadjuvantly treated patients in the TransNEO cohort that they predicted pathological complete response with odds ratios of 1.95 (95% CI 1.22-3.03, p=0.005) and 2.32 (95% CI 1.43-3.77, p=0.001), and in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas showed overall survival hazard ratios of 0.79 (95% CI 0.63-1.00, p=0.05) and 0.80 (95% CI 0.67-0.97, p=0.02), with correlation to gold standard clinical assessment of 0.59-0.69 and no substantial difference from gold standard assessment in predicting pathological complete response (AUC 0.60-0.
The latest research from Google

Retrieve-for-Train: Compiling Query Fan-Out Offline with RL to Bypass Inference Latency via a Diffusion Retriever

The work proposes Retrieve-for-Train, which first trains a fan-out language model with offline reinforcement learning (built on Gemma3-4B and Qwen3-4B, emitting 10 sub-queries per prompt) under a composite reward of groundedness, Vendi-Score diversity, and alignment that scores the whole result set, then distills that behavior into a 53.9M-parameter diffusion retriever that generates the complete target set in one non-autoregressive parallel pass in continuous embedding space, outperforming single-query search, zero-shot expansion, and a Best-of-N baseline on open-ended abstract retrieval and weakly supervised compositional retrieval while achieving a 12 to 20 speedup over autoregressive approaches.
Nature News

Turning a paper into a conversational agent: reading the Paper2Agent report

According to a Nature news report, Paper2Agent reads a paper's main text, code and data sets, deposits them on an MCP server, and has a team of AI agents autonomously build tools that apply the paper's methods, producing a paper-specific agent that can be questioned in plain language as a 'virtual corresponding author'; the report says it created an agent for the AlphaGenome paper in about 45 minutes at US$14 of computing cost, that the agent answered genetics questions with near-perfect accuracy and outscored other top biomedical AI agents including Biomni, and that it was used to re-examine which causal gene explains a single-letter DNA change linked to 'bad' cholesterol, arriving at a different gene from the one pinpointed in the original paper.
Nature News

Human Brain Tissue Transplanted into Mice: The Most Extensive Brain-Region Chimeric Model to Date

By genetically engineering mice so that the precursor cells that would have formed the cerebral cortex do not survive, thereby emptying part of the brain cavity, and then implanting human brain tissue into newborn mice, the study reports that the graft expanded nearly fivefold within two to three months, filled more than 90% of the vacant space, sent projections deep into the rodents' spinal cords, and developed specialized neurons including cells similar to von Economo neurons, while behavioral tests showed no enhancement of the rodents' intellect.
Nature News

Cultured Asgard Archaea and a Virus-Host System: Live Clues to Eukaryogenesis

This evidence bundle, comprising a Nature news feature and two bioRxiv preprints, reports progress in culturing Asgard archaea (Promethearchaeota): live-cell microscopy showed that cells of the Loki and Hodarchaea lineages drastically change shape on a minute timescale, extend and retract protrusions at 1.5 to 5.3 micrometres per minute, and crawl on glass surfaces, with actin inhibitors arresting these dynamics; separately, an Asgard archaeal virus infecting a novel strain of Ca. Lokiarchaeum ossiferum B36 was cultured for the first time, with a 16 kbp integrated provirus able to excise and replicate independently to form virus particles, placed in a new family Fylgjaviridae, while the host carries Septu, Wadjet and type II CBASS antiviral defence systems.
Nature News

Map of brain 'microproteins' could offer new clues to Alzheimer's disease

Combining mass spectrometry, RNA sequencing and ribosomal profiling on 608 post-mortem dorsolateral prefrontal cortex samples from individuals with and without Alzheimer's disease, the study identified 4,321 microproteins (3,217 of them not previously characterized in the UniProtKB/Swiss-Prot standard human protein catalogue), applied a deep-learning model to rank the mass-spectrometry identification confidence of 3,001 of them with 1,067 rated high confidence, and found dozens of microproteins with altered expression in people with Alzheimer's disease, producing what is described as the largest atlas of microproteins in Alzheimer's disease made so far and making the dataset publicly available.