Skip to main content

Medicine & Health

469 items

  1. medRxiv

    Beyond word error rate: clinical risk as the necessary standard for ambient AI scribe evaluation: evidence from 77 global languages

    This study constructed a multilingual clinical dictation corpus (five clinical dictation scripts spanning a complexity gradient, translated into 99 languages, rendered to synthetic speech under three acoustic conditions, and transcribed by a production ambient scribe), computed six frequency metrics, and had three independent large language model raters assess clinically meaningful error patterns using a Severity x Likelihood framework; across 59,819 genuine transcription-error occurrences, 58,329 (97.5%) were LOW risk and 251 (0.42%) CRITICAL or HIGH, none of the six frequency metrics showed a detectable association with serious clinical risk (absolute Spearman rho < 0.16), a Severity x Likelihood sum remained strongly correlated with WER (rho=0.
  2. Nature Medicine

    Prospective evidence for conversational medical AI: hard, but non-negotiable

    This Nature Medicine comment argues that trust in clinical AI cannot be benchmarked into existence but must be earned through rigorous prospective studies in real-world clinical settings, noting that the hardest lessons often concern the humans and systems around the AI rather than the technology itself.
  3. npj Digital Medicine

    Multi-Agent Collaboration as a Complementary Architecture for AI-Generated Medical Examination Items

    Responding to Qian et al.'s finding that a single LLM can generate acceptable knowledge-based questions but struggles with higher-order reasoning, this work proposes MAID, a multi-agent architecture that decomposes item development into specialized agents for drafting, critique, and iterative adversarial refinement, and in a blinded paired-comparison evaluation of items aligned with China's National Medical Licensing Examination standards, 14 experts from 7 medical disciplines compared 25 matched MCQ pairs yielding 350 item-level observations under a two-alternative forced-choice design, with multi-agent outputs receiving 57.7% of expert preferences (95% CI [52.5%, 62.8%]) versus 42.3% for the single-model baseline (95% CI [37.2%, 47.5%]; χ²(1) = 8.33, p = 0.
  4. medRxiv

    The Multimodal Anonymizer: a fully local multi-agent AI system for medical data deidentification

    The study developed and evaluated the Multimodal Anonymizer, a modular, locally deployable multi-agent framework integrating multimodal large language models, task-specific neural networks, and rule-based transformations; on benchmarks spanning text, tables, PDFs, imaging, metadata, filenames, audio, handwriting, and 3D imaging, its best local configuration (orchestrator Qwen3-VL-235B-A22B-Thinking) achieved 98.80% per-patient deidentification sensitivity (95%-CI 97.20; 100) and 99.60% critical clinical preservation (95%-CI 98.80; 100), reached 100% sensitivity and critical preservation on 250 local Charité partograms, and outperformed established tools across most modalities.
  5. Drug Discovery Today

    From implicit prioritization to auditable decisions in natural product drug discovery

    The article proposes a framework for natural product drug discovery that translates evidence into auditable experimental actions through a four-tier architecture (source authentication, reproducible chemical fingerprinting, orthogonal prioritization, definitive characterization) with documented decision gates, and an 'evidence ladder' that separates molecular-identification confidence from chemical novelty, biological novelty and translational relevance, addressing three recurring gaps: incomplete integration of sample metadata, conflation of analytical detection with molecular novelty, and lack of explicit decision thresholds.
  6. Advances in wound care

    Admission Red Cell Distribution Width-Albumin Ratio and 1-Year Mortality in Chronic-Wound Inpatients: Prediction-Model Development, Internal Validation, and Exploratory Temporal Evaluation

    In a single-center retrospective cohort of 584 adults first admitted for a chronic wound between 2021 and 2024, a two-variable model combining the admission-day red cell distribution width-albumin ratio (RAR) with the age-adjusted Charlson Comorbidity Index (ACCI) predicted 1-year all-cause mortality with an optimism-corrected C-statistic of 0.855 and good calibration, stratifying patients into low-, intermediate-, and high-risk tiers (observed mortality 1.3%, 9.0%, and 36.8%), while discrimination was similar but calibration imprecise in an exploratory temporal evaluation of a later same-center cohort of 124 patients.
  7. npj Digital Medicine

    BenchECG and xECG: A Standardized Benchmark and Baseline for ECG Foundation Models

    This work introduces BenchECG, a standardized benchmark spanning eight public ECG datasets, 421,171 patients, 1,674,704 recordings, and ten tasks (classification, segmentation, detection, regression, survival analysis), used to evaluate public foundation models such as ST-MEM, ECG-JEPA, and ECGFounder; it also proposes xECG, a bidirectional xLSTM model pretrained with SimDINOv2 self-supervised learning, which achieves the best BenchECG score of 0.868±0.0030 (mean rank 1.50 under finetuning and 1.20 under linear probing), is the only public model to perform strongly across all datasets and task types, and leads on long-context tasks (sleep apnea AUROC 0.932±0.014; MIT-BIH arrhythmia F1 0.677±0.025) and computational efficiency (about 10x less time and about 7x less memory on PTB-XL).
  8. Journal of Chemical Information and Modeling

    AI-Driven Drug-Target Interaction Prediction: From Data Representation to Model Design — A Systematic Review

    This review systematically surveys AI-driven drug-target interaction (DTI) prediction, starting from classical molecular binding theories (lock-and-key, induced fit, conformational selection) and summarizing task settings such as binary interaction classification, binding affinity regression, and multitask prediction with uncertainty assessment; it organizes multimodal representations for drugs and target proteins (molecular sequences, graph structures, 3D conformations, physicochemical properties, biological perturbation profiles, protein sequences and structures, biomedical knowledge networks) together with interaction labels and auxiliary biomedical data, compares representative approaches across orthogonal dimensions including input representation, encoder architecture, interaction-mod
  9. Pharmaceutical Medicine

    From Compliance to Strategic Partner: The Transformation of Regulatory Affairs in AstraZeneca Local Affiliates

    This article describes the transformation of AstraZeneca's local Marketing Companies Regulatory Affairs (MCRA) teams in Europe from a predominantly compliance-driven support function into a strategic partner in drug development and patient access, organized around three pillars (Launch Excellence, External Engagement and Advocacy, and Digitalisation and Process Optimisation) and a "One-Team" local-global regulatory culture, and proposes a contribution-based measurement framework to monitor progress.
  10. Current Neurology and Neuroscience Reports

    Advances in Multimodality Monitoring in Traumatic Brain Injury

    This review evaluates the current state of multimodal monitoring (MMM) technologies in traumatic brain injury (TBI), their clinical applications, and their emerging role in improving TBI management, focusing on real-time physiological data streams such as intracranial pressure, brain tissue oxygenation, arterial blood pressure, and electroencephalography; it suggests that integrating physiological variables with artificial intelligence and advanced analytics may enable earlier detection and prediction of adverse events such as seizures, brain hypoxic events, and excursions of cerebral hypertension, potentially reducing variability in TBI outcomes associated with secondary insults, while the optimal combination of monitoring modalities, interpretation of complex data streams, larger harmoni

Page 43 · showing 10