Skip to main content

Medicine & Health

466 items

  1. arXiv

    OPGAgent uses a multi-tool agent with consensus to read panoramic dental X-rays, reaching 42.3% exact-match F1 on its OPG-Bench while holding false positives to 4.89 per case

    The work proposes OPGAgent, a multi-tool dental agent planned by GPT-5.2 under the ReAct paradigm that orchestrates hierarchical evidence gathering, a specialized toolbox, and a consensus subagent, together with OPG-Bench, a structured-report protocol built on (Location, Field, Value) triples derived from real clinical reports; on OPG-Bench, comprising 1,009 anonymized OPGs and 5,219 VQA pairs, OPGAgent reaches 42.3% exact-match F1, a 49.7% aggregate score, 43.1% precision and 4.89 false positives per case, and leads MMOral-OPG at 62.53% accuracy.
  2. arXiv

    ConRad fine-tunes a medical vision-language model with GRPO and a logarithmic scoring rule, cutting report-level ECE from 0.642 to 0.106 and improving zero-shot IU-Xray ECE from 0.310 to 0.034

    The work introduces ConRad, a reinforcement learning framework that fine-tunes medical large vision-language models with the GRPO algorithm and a logarithmic scoring-rule reward so they emit verbalized confidence alongside radiology reports at report or sentence level, reducing report-level ECE to 0.106 and sentence-level ECE to 0.239 on MIMIC-CXR and improving zero-shot IU-Xray ECE from 0.310 to 0.034 while keeping report quality (GREEN) stable.
  3. arXiv

    MPFlow guides a rectified-flow prior with auxiliary MRI at inference, matching diffusion baselines at 20% of sampling steps and cutting tumor-hallucination Dice by 15%

    The work proposes MPFlow, a zero-shot multi-modal MRI reconstruction framework built on rectified flow that uses a self-supervised pretraining strategy, PAMRI, to learn shared cross-modal representations and jointly guides the unconditional prior with data consistency and cross-modal feature alignment at inference; on HCP T2 4x super-resolution and BraTS FLAIR 8x k-space reconstruction it matches diffusion baselines in image quality using only 20% of the sampling steps while improving tumor segmentation Dice by 15% and reducing the SHAFE hallucination score by 26%.
  4. arXiv

    RPG-SAM pairs reliability-weighted prototypes with geometric adaptive thresholds to lift training-free one-shot polyp segmentation by 5.56% mIoU on Kvasir

    The work proposes RPG-SAM, a SAM2-based training-free one-shot polyp segmentation framework that uses Reliability-Weighted Prototype Mining (RWPM) to weight support foreground prototypes by a contrast factor and a reverse purity factor while using background prototypes as negative anchors for noise suppression, Geometric Adaptive Selection (GAS) to pick binarization thresholds dynamically from morphological solidity and scale consensus, and a Prior-guided Iterative Refinement (PIR) loop to polish boundaries, reaching 78.65% mIoU and 85.65% mDice on Kvasir, surpassing ProtoSAM by 5.56% and 4.11% respectively, with comparisons also reported on three-center PolypGen, CVC-ClinicDB, and CVC-ColonDB.
  5. arXiv

    Fact-Flow uses LLM-bootstrapped multi-label fact guidance to raise factual accuracy of MLLM medical reports on tuberculosis and ophthalmology data

    The work introduces Fact-Flow, a framework that decouples visual fact identification from report generation: an LLM automatically extracts and merges clinical finding labels from training reports (7 labels for tuberculosis, 42 for ophthalmology), a multi-label classifier predicts findings, and the predicted labels are serialized into a prompt to guide an MLLM; on the tuberculosis chest X-ray dataset MedGemma + Fact-Flow reaches RadFact F1 0.3055 versus 0.2266 for MedGemma alone, and on the ophthalmology dataset Qwen2.5-VL + Fact-Flow is best on most NLG metrics.
  6. Frontiers in Pharmacology

    Review proposes a microbiome-aware oral formulation framework: excipient risk classification, nanocarrier design rules, and a tiered testing roadmap

    This review integrates evidence on microbial drug metabolism, excipient–microbiome interactions, nanocarrier–microbiome interfaces, microbiota-responsive release mechanisms, and experimental models to propose an authors' evidence-informed framework for oral formulation design, comprising an excipient–microbiome risk classification, nanocarrier design rules, microbiota-responsive delivery decision logic, and a tiered testing roadmap, illustrated by clinically relevant examples such as digoxin inactivation by Eggerthella lenta, bacterial levodopa metabolism, and microbial beta-glucuronidase-mediated irinotecan toxicity.
  7. arXiv

    ClinCoT pushes preference optimization from answer-level correction down to lesion-region reasoning: it beats MMedPO and other baselines on most metrics across SLAKE, VQA-RAD and IU-Xray, with more consistent gains after SFT initialization

    The work proposes ClinCoT, a clinical-aware visual chain-of-thought framework: it uses hypotheses-driven region proposals (disease-conditioned activation maps from a clinical-aware tool such as MedKLIP, thresholded into candidate regions), has the target Med-VLM generate multiple region-conditioned reasoning chains by jointly processing the original image and each candidate region, scores them with two Med-LLM evaluators that combine current-step score with next-step influence and penalize evaluator disagreement, and trains with a score-margin-aware DPO plus iterative regeneration of preference data; across SLAKE, VQA-RAD and IU-Xray it outperforms DPO, Self-Rewarding, STLLaVA-Med, POVID, SIMA, FiSAO and MMedPO on most metrics, is strongest on report generation, falls below MMedPO on VQA-R
  8. arXiv

    AbSteering steers general-purpose VideoLMs with abnormality-centric chain-of-thought and DPO to generate HRCT reports, surpassing large-scale CT-specific foundation models on fine-grained clinical metrics while improving detection sensitivity and reducing hallucinations.

    The work presents AbSteering, a two-stage framework combining abnormality-centric chain-of-thought training with a Direct Preference Optimization objective for fine-grained abnormality discrimination, to adapt general-purpose VideoLMs to high-resolution CT report generation, and curates the CT-RATE-AB dataset; results show that general-purpose VideoLMs transfer effectively to 3D medical imaging under limited data, achieving state-of-the-art performance on fine-grained clinical efficacy metrics, with superior detection sensitivity over domain-specific CT foundation models pretrained on large-scale CTs while mitigating hallucinations.
  9. International Journal of Computer Assisted Radiology and Surgery

    NeuralShift predicts brain shift in temporal lobe resection from preoperative MRI alone, reaching Dice 0.97 and landmark TRE as low as 1.12 mm

    The study introduces NeuralShift, a U-Net-based model that takes only preoperative MRI plus a hemisphere indicator encoding resection laterality and predicts a dense displacement field mapping preoperative to intraoperative MRI, the intraoperative brain mask, and its signed distance function; evaluated on 98 paired preoperative and intraoperative T1-weighted MRI scans from epilepsy patients undergoing temporal lobe resection with 9-fold cross-validation, it achieved a Dice of 0.97±0.01 between predicted and intraoperative masks (versus 0.92±0.01 for the preoperative mask) and reduced landmark TRE on the resection side and midline from about 4.58 mm to about 2.96 mm (left) and from about 4.41 mm to about 2.89 mm (right), with a minimum of 1.12 mm.
  10. arXiv

    IOSVLM diagnoses multiple dental diseases directly from native 3D intraoral scan point clouds, reaching 77.23% macro accuracy, 9.58 points above Gemini 3 Pro

    The authors present IOSVLM, an end-to-end 3D vision-language model that represents intraoral scans (IOS) as point clouds and follows a 3D encoder-projector-LLM design for unified diagnosis and generative VQA, together with IOSVQA, a dataset of 19,002 cases and 249,055 VQA pairs over 23 oral diseases and heterogeneous scan types, and with a geometry-to-chromatic proxy plus two-stage curriculum training it reaches 77.23% macro accuracy and 50.39% macro F1 on IOSVQA, outperforming baselines including GPT-5 and Gemini 3 Pro.

Page 19 · showing 10