Skip to main content

Medicine & Health

467 items

  1. Studies in health technology and informatics

    Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research

    This study had five current large language models (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) reproduce three statistical tasks from a published clinical workflow under identical datasets and standardised prompts—descriptive table generation, Kaplan–Meier survival analysis, and Cox proportional hazards modelling—and compared their outputs with analyses by two experts trained in mathematical statistics on grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code, finding that all five models produced correct descriptive statistics once dataset variables were specified explicitly, that two models failed the initial descriptive benchmark because of variable-name ambiguity but recovered after prompt clarification, that all models
  2. Studies in health technology and informatics

    Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes

    This pilot study anonymized 35 German doctor's notes from five patients, built one pipeline for medication extraction and mapping and two for diagnoses (one RAG-based and one agentic AI), ran them with three open-weight LLMs on a local GPU-PC, and found that medication name extraction reached an F1 of 0.95 and medication mapping 0.78, while diagnosis coding did not exceed an F1 of 0.12 and broad-category mapping reached 0.18, leading the authors to conclude that LLMs are suitable for medication information extraction for research databases but that current state-of-the-art open-weight models are not accurate enough for a clinical setting where patient treatment would depend on LLM performance.
  3. IEEE transactions on medical imaging

    SurgDepth: Test-Time Adaptation of Depth Foundation Models for Surgical Scene Understanding

    The work proposes SurgDepth, a test-time adaptation framework that adapts natural-image-pretrained depth foundation models to surgical endoscopy without any labeled surgical data, combining two self-supervised signals (stereo photometric consistency and flip equivariance, the latter requiring only single images) with selective decoder adaptation and two tuning-free reliability mechanisms (progress-aware anchor regularization and a structural-trust safeguard for poorly-illuminated sequences); the authors report AbsRel reductions of up to 63% with stereo pairs and 42% without, improvement in four of five cross-domain evaluation settings, generalization across seven foundation models, and the lowest AbsRel (0.
  4. World journal of urology

    Multimodal Large Language Models for Bladder Tumor Detection in Cystoscopy: A Retrospective Benchmarking Study

    This retrospective study analyzed 1,754 labeled public cystoscopy images to test Direct, Book-based, and Optimized prompts across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano for benign-versus-malignant classification, finding that GPT-5 and GPT-5-Mini with the optimized prompt reached accuracies of 86.7% and 89.2%, and that GPT-5 with the optimized prompt achieved 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity in high-confidence triage at 62.0% image coverage, while prompt engineering improved calibration and triage without statistically significant performance gains.
  5. Molecular Biomedicine

    Decoding neuro-tumor interactions in pancreatic cancer: mechanisms, immunosuppressive networks and therapeutic opportunities

    This review systematically examines the molecular mechanisms of perineural invasion (PNI) in pancreatic ductal adenocarcinoma, proposes a unified four-stage model spanning mutual chemotaxis between tumors and nerves, adhesion and invasion at the tumor-nerve interface, extracellular matrix remodeling, and neural plasticity alterations, defines the perineural invasion microenvironment as a neuro-immune privileged sanctuary, and summarizes therapeutic strategies targeting the neuro-immune-tumor axis, ongoing clinical trials, and the applications of multi-omics and artificial intelligence in PNI diagnosis, mechanistic discovery, and therapeutic optimization.
  6. Journal of medical systems

    Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction & Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes

    This report develops CLASS, a Python-based modular large language model pipeline that runs within a secure institutional environment and combines expert-curated concept lists, a task-specific prompt suite, and a schema-constrained output format to extract structured data from clinical notes, and evaluates it exploratorily on a single-center retrospective corpus of pediatric esophageal airway treatment surgery (EATS) operative notes: observed concordance with surgeon adjudication on the 20 longest notes (3,960 note-procedure pairs) was high (F1 0.9967), while CLASS proposed 28 candidate procedure variants or additions, 18 (64.
  7. Journal of the American Medical Informatics Association : JAMIA

    DiagnosticXchange: An Open-Source Framework for Evaluating Safety, Efficiency, and Diagnostic Reasoning in Clinical AI Systems

    This study developed and validated DiagnosticXchange, an open-source clinical simulation framework in which AI systems diagnose cases by ordering tests, requesting imaging, and performing procedures, with each action mapped to CPT codes capturing cost, time, work relative value units, and invasiveness; using 8 large language models on 216 peer-reviewed cases across 19 specialties (1728 sessions), it found that three systems with near-identical accuracy (93.5%–94.0%) differed significantly in cost (P < .001; 1.75-fold between the most and least expensive) and 2.
  8. PLOS digital health

    Utilization of a HIPAA-compliant large language model chatbot in an academic pediatric medical center

    This mixed-methods case study analyzed 14 months of utilization of "InternalGPT," a HIPAA-compliant LLM chatbot at an academic pediatric medical center, finding that 2,149 of approximately 15,788 employees (13.6%) requested access, 52.8% recorded at least one token use, the top 20% of users consumed 69.4% of tokens, and among 461 sustained users, 92 self-report survey respondents indicated a mean 30% productivity gain corresponding to an exploratory perceived productivity value of $6.3M to $18.9M under varying extrapolation assumptions.
  9. Studies in health technology and informatics

    Medical Concept Normalization of German Clinical Expressions to SNOMED CT: Domain Embedding Retrieval with LLM Reranking Outperforms LLM-Only

    This study investigates medical concept normalization of short German clinical expressions to SNOMED CT by comparing a direct GPT-5.4 LLM-only approach with a hybrid approach combining medBERT.de bi-encoder embedding retrieval and RAG-based LLM reranking, finding that the LLM-only baseline achieves Recall@1 of 0.235, Recall@3 of 0.297, and Recall@5 of 0.303, while embedding-based retrieval reaches Recall@1 of 0.681, Recall@3 of 0.783, and Recall@5 of 0.812, and adding RAG reranking further improves Recall@1 to 0.771 with Recall@3 and Recall@5 at 0.809 and 0.812.
  10. bioRxiv

    3D Spatial Interactomics Maps the Dynamics of NF-κB Multiprotein Signalosomes in Single Cells

    This work introduces an intelligent sequential proximity ligation assay (iseqPLA) read out by spinning disk confocal microscopy and 3D reconstruction to profile endogenous NF-κB protein-protein interactions inside single cells, treating clusters of co-localized puncta as a measure of supercomplex spatial organization, and tracks supercomplex dissociation, p65 nuclear translocation, and negative-feedback engagement across cytokine time courses in NIH-3T3 mouse fibroblasts, cystic fibrosis (CF) patient-derived macrophage co-cultures with IMR-90 human fibroblasts, and an independent set of healthy- and CF-donor monocyte-fibroblast co-cultures, reporting three findings: 3D volumetric quantification reduces the variance in nuclear-to-cytoplasmic ratio measurements relative to 2D projections, th

Page 30 · showing 10