Skip to main content

Medicine & Health

463 items

  1. arXiv

    BiasRecon hits top PSNR in cross-anatomy, cross-center and cross-modality MRI reconstruction with fewer than 100 tunable parameters

    The work proposes BiasRecon, a bias-calibrated adaptation framework grounded in a minimal-intervention principle, which alternates frequency-guided prior calibration (learnable scalars α and β modulating low- and high-frequency components at each U-Net skip layer, adding only 2L parameters), score-based denoising, and adaptive regularization that uses Stein's Unbiased Risk Estimator (SURE) to tune the regularization parameter γ; trained on FastMRI knee (973 volumes) and evaluated without retraining on FastMRI-Brain (cross-anatomy), Stanford Knee (cross-center) and CMRxRecon (combined shift) under 8× Gaussian 1D undersampling, it improves PSNR over DDS by +1.19 dB, +2.06 dB and +1.60 dB respectively, and is best in 10 of 12 sampling/acceleration settings.
  2. arXiv

    UltraStar recasts echocardiography probe navigation from path regression to anchor-based global localization, beating baselines and scaling better with longer inputs on over 1.31 million samples

    The work proposes UltraStar, which reformulates echocardiography probe navigation from path regression to anchor-based global localization: a Star Graph treats historical keyframes as spatial anchors connected directly to the current view to explicitly model geometric constraints, and a semantic-aware sampling strategy actively selects representative landmarks from massive history logs to reduce redundancy for accurate anchoring; experiments on a dataset with over 1.31 million samples show it outperforms baselines and scales better with longer input lengths, indicating a more effective topology for history modeling under noisy exploration.
  3. arXiv

    K-MaT aligns prompt manifolds via optimal transport to transfer medical VLMs to low-end modalities without low-end training images

    K-MaT is a prompt-learning framework that factorizes prompts, anchors them to clinical text descriptions, and aligns the low-end prompt manifold to the visually-grounded high-end space using Fused Gromov-Wasserstein optimal transport, transferring decision structures to low-end modalities without requiring low-end training images; across four cross-modal benchmarks including dermoscopy, mammography to ultrasound, and CT to chest X-ray it achieves state-of-the-art results, raising the average harmonic mean of accuracy to 44.1% from BiomedCoOp's 42.0% with macro-F1 of 36.2%, and on the challenging breast imaging task it mitigates the catastrophic forgetting seen in CoOp, which drops to 27.0% accuracy on the low-end.
  4. arXiv

    OncoAgent turns esophageal radiotherapy guideline text into 3D target volumes zero-shot, reaching CTV Dice 0.842 with no significant difference from fully supervised nnU-Net

    The study introduces OncoAgent, a guideline-aware AI agent framework in which a large language model parses free-text clinical guidelines into an executable tool-call plan, generating three-dimensional target volumes without any expert-annotated training data; on planning CT from 40 mid-thoracic esophageal cancer patients (32 for training, 8 for testing), it achieved zero-shot CTV Dice of 0.842 and PTV Dice of 0.880, with no statistically significant difference from the fully supervised nnU-Net (GTV Prior) at 0.862 and 0.893 on primary metrics, and it was rated higher than that supervised baseline by blinded physicians on guideline compliance, modification effort, and clinical acceptability.
  5. arXiv

    IntraStyler uses a contrastively trained style encoder to synthesize T2 MRI styles without predefined sub-domains, raising Extra-VS Dice from 0.61 to 0.81 with zero failures on CrossMoDA

    The work proposes IntraStyler, a 3D unpaired image translation method that requires no predefined sub-domains, using a contrastively trained style encoder to extract anatomy-disentangled style embeddings from each target-domain image and conditioning a QS-Attn synthesis network via dynamic instance normalization; on the CrossMoDA benchmark (226 labeled ceT1, 295 unlabeled T2, 96-image T2 test set) it automatically discovers 7 style references instead of the 3 used by sub-domain methods, and downstream nnU-Net segmentation achieves the best Dice and ASSD on Intra-VS, Extra-VS, and Cochlea, with Extra-VS Dice of 0.81 and zero failures.
  6. arXiv

    MedQ-Engine lifts an 8B medical image quality model past GPT-4o by over 13% and cuts the gap to human experts to 4.34% using 10K annotations

    The work proposes MedQ-Engine, a closed-loop data engine that iterates through evaluate-explore-evolve phases: it clusters model failures on a development set into failure prototypes, uses the visual component of those prototypes as retrieval anchors over a roughly one-million-image pool spanning five modalities, and combines progressive human-in-the-loop annotation with entropy-guided routing and quality-assured fine-tuning, so that with only 10K annotations an 8B-parameter model surpasses GPT-4o by over 13% on medical image quality assessment and narrows the gap to human experts to 4.34%, with more than 4x sample efficiency over random sampling.
  7. arXiv

    Splitting prompt dependence into prompt ambiguity and local sensitivity yields two low-correlated metrics that both correlate negatively with gynecological pelvic MRI segmentation quality

    The work introduces the first framework that explicitly disentangles prompt dependence into prompt ambiguity (inter-user variability) and local sensitivity (interaction imprecision): a mixture density network models the image-conditioned distribution of plausible prompts and quantifies ambiguity by the trace of its covariance, while a stability margin captures the minimal perturbation under measurement noise needed to induce a mask change; evaluated on two female pelvic T2-weighted MRI datasets, UT-EndoMRI (81 patients with endometriosis, uterus segmentation) and MOGaMBO (94 patients with locally advanced cervical cancer, bladder segmentation), with MobileSAM and MedSAM, the two metrics correlate significantly and negatively with Dice, correlate little with each other, and conditional samp
  8. arXiv

    An RL-fine-tuned small model lets a robot follow ultrasound guidelines to scan gallbladder, spine, and kidney autonomously

    The work proposes an autonomous robotic ultrasound framework driven by an LLM agent that retrieves guideline steps from scanning handbooks, reasons over current observations and scanning state, and dynamically invokes tools for trajectory planning, robot execution, contact adjustment, voice guidance, and trajectory refinement, with reinforcement-learning (PPO) fine-tuning to improve reasoning quality and the correctness of tool selection and parameterization; validated verbally on 10 unseen ultrasound scanning guidelines, fine-tuning raised step-wise accuracy from 0.6512 to 0.8973 and overall success rate from 0.5384 to 0.
  9. arXiv

    SEA-PEFT raises mean Dice by 2.4–2.8 points in 1/5/10-shot 3D medical segmentation while training under 1% of parameters

    The authors propose SEA-PEFT, which treats adapter configuration as an online allocation problem solved during fine-tuning through a search–audit–allocate loop that trains active adapters, estimates each adapter's Dice utility by momentarily toggling it off, and reselects the active set under a parameter budget with a greedy knapsack allocator, stabilized by EMA+IQR smoothing and a finite-state machine; on TotalSegmentator and FLARE'22 it improves mean Dice by 2.4–2.8 points over the strongest fixed-topology PEFT baselines across 1/5/10-shot settings while training under 1% of parameters.
  10. arXiv

    Treating electrode coordinates as learnable parameters: percept-aware optimization on folded cortex improves reconstruction fidelity while eliminating vascular safety-margin violations

    The study presents a percept-aware surgical planning framework for cortical visual prostheses that treats 3D electrode coordinates as learnable parameters and optimizes them end-to-end through a differentiable forward model of prosthetic vision on FreeSurfer fsaverage folded cortical geometry, minimizing task-level perceptual error subject to vascular avoidance and gray matter feasibility constraints; on simulated MNIST reading and CIFAR-10 natural image tasks it consistently improved reconstruction fidelity over visual field tiling and visual field coverage baselines (median MSE reductions of up to 67.7% and 33.4% on MNIST, with downstream classification accuracy gains of 62.6% and 22.4%), eliminated all 300 µm safety-margin violations with only 1.7% (MNIST) and 4.

Page 15 · showing 10