Medicine & Health
467 items
Embedding-Based Evaluation and Benchmarking Framework for Optimizing LLM Post-Processing of Medical Transcriptions from a Local Whisper System
This study proposes a reference-free evaluation framework that combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation to benchmark LLM post-processing models for clinical transcripts produced by a locally deployed Whisper system, applying it to 26 German physician-patient conversations in which four LLMs generated 416 summaries, and finding that GPT-OSS-120B and MedGemma-27B achieved the highest similarity scores across most embedding models, that cosine similarities across repeated runs typically exceeded 0.90, and that embedding-based similarity signals only partially agreed with human judgments of summary quality.
Deterministic and stochastic interventions in reducing drug-drug interactions in inappropriate prescribing: A systematic review
This systematic review searched PubMed, Scopus, ScienceDirect, and IEEE Xplore following PRISMA 2020 and the SPIDER framework, included 10 primary studies of computational and clinical decision support interventions aimed at reducing drug-drug interactions or inappropriate prescribing, synthesized them narratively across deterministic, ontological, and stochastic/generative categories, and assessed risk of bias with PROBAST+AI, finding that earlier deterministic systems showed modest improvements in prescribing process measures with inconsistent links to patient-level outcomes, that recent stochastic and generative models reported strong internal performance metrics, and that AI-driven studies carried a consistently high risk of bias in the analysis domain driven mainly by limited external
Artificial intelligence-assisted lead optimization in drug discovery: bridging computational advances and translational challenges
This review systematically surveys the current landscape of artificial intelligence and machine learning in lead optimization, covering advances in graph neural networks, transformer architectures, diffusion models, and chemical foundation models for molecular design and property prediction, and discusses emerging concepts including data-centric AI, uncertainty quantification, trustworthy AI, the AI optimization paradox, and the shift from molecular prediction toward scientific decision-making, concluding that despite increasing industrial adoption, AI remains dependent on high-quality experimental data, model generalizability, and rigorous experimental validation, and that future progress will depend less on increasingly sophisticated algorithms than on trustworthy AI systems that improve
M3: Conversational LLMs simplify secure clinical data access, understanding, and analysis
The study introduces and evaluates M3, a Python server system built on the Model Context Protocol (MCP) that lets researchers query the MIMIC-IV critical care database in natural language; on 100 answerable and 100 unanswerable EHRSQL 2024 questions, Claude Sonnet 4 reached 94% accuracy and the open-weights gpt-oss-20B, deployable locally on consumer hardware, reached 93% (McNemar p=1.00, no detectable difference), while gpt-oss-20B correctly abstained on 69% of unanswerable questions, with errors mainly reflecting question ambiguity and mismatched schema mappings.
An expert-level generalist AI for abdominal CT diagnosis: RADAR
This work developed RADAR, a generalist vision-language model trained on more than 400,000 contrast-enhanced abdominal CT examinations and 15 million anatomy-wise image-text pairs, learning directly from clinical reports without manual annotation, achieving high diagnostic performance and robust generalization across internal and external evaluations for 18 anatomical structures and 146 imaging findings, and increasing the diagnostic sensitivity of 26 radiologists by ~10% in a reader study.
Evaluating Large Language Models for Lay Summaries of Radiology Reports Using Tailored Prompting Strategies and Mixed-Method Assessment
Using 100 radiology reports from the BioNLP 2023 report summarization dataset, this study had five large language models generate patient-facing lay summaries under individually tailored prompting styles selected in pilot work (few-shot for GPT-4; generated knowledge for GPT-4o mini, Gemini 1.5 Pro and Gemini 1.5 Flash; zero-shot for Llama 3.1), then evaluated them through a mixed framework of two radiology fellows, two large reasoning models (Gemini 2.5 Pro and GPT-oss-120b), and readability metrics, finding that Gemini 1.5 Flash and Pro with generated knowledge ranked highest for actionable content, minimal-supervision usability and readability, that GPT-4 with few-shot achieved the highest human-rated accuracy (98%), and that expert and large-reasoning-model ratings aligned closely.
Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study
This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
Segmentation of placental tissue using immunofluorescent staining and an artificial intelligence-based analysis workflow
The study developed a HALO AI-based placental tissue classifier that uses PLAP and DAPI staining to segment sections into villous core, villous trophoblast (VT), and intervillous space (IVS), validated area and intensity measurement in classified regions with SDC-1 and vimentin staining, and increased the area of staining analyzed to 245 times that of a single field of view on whole slide scanning images.
Metagenomic Deep Sequencing Identifies Gene Mutations Associated with Chemotherapeutic Resistance in Vitreoretinal Lymphoma
In 49 patients with vitreoretinal lymphoma (VRL) confirmed by cytopathology, immunohistochemistry, flow cytometry, and/or MYD88 PCR, host-genome metagenomic deep sequencing (MDS) of intraocular specimens cross-referenced against the Catalogue of Somatic Mutations in Cancer identified eight gene mutations associated with chemotherapeutic resistance in six specimens from four patients, including methotrexate-resistance-associated mutations in four specimens from three patients, and in one patient serial sampling at initial vitrectomy and two subsequent recurrences revealed distinct resistance-associated mutations at each time point; that patient died despite multi-agent therapy including rituximab, consolidation regimens, and lenalidomide, whereas the other three patients remained in long-te
Page 32 · showing 10