Skip to main content

Medicine & Health

466 items

  1. arXiv

    DP-RGMI splits differential privacy's performance loss on 594,000 chest X-rays into encoder geometry and task-head utilization

    The authors introduce DP-RGMI, a framework that treats differentially private training as a structured transformation of representation space and decomposes performance degradation into representation displacement, spectral effective dimension, and a utilization gap defined as the difference between linear-probe and end-to-end AUROC; across more than 594,000 chest X-rays from four datasets and three pretrained initializations (ImageNet, DINOv3, MIMIC-CXR), with PadChest as the primary dataset (110,525 frontal images, 22,045 test images), they find that strong privacy consistently leaves linear separability largely preserved while a utilization gap persists (G = 8.0 for ImageNet at ε = 1.0, 3.4 for MIMIC at ε = 0.7, 6.1 for DINOv3 at ε = 0.
  2. arXiv

    HD-TTA chooses between competing 'compact or inflate' hypotheses, cutting HD95 by about 6.4 mm and raising precision by over 4% in cross-domain brain tumor segmentation

    The work proposes Hypothesis-Driven Test-Time Adaptation (HD-TTA): with a frozen nnU-Net v2 backbone and optimization only over test-sample logits, a Gatekeeper first decides whether a case needs refinement, two competing geometric hypotheses (compact denoising vs. diffuse recovery) are generated in parallel, and a representation-guided selector picks the safest output using intrinsic texture consistency; trained on BraTS 2023 GLI and evaluated with strictly fixed hyperparameters on unseen pediatric (PED) and meningioma (MEN) target domains, HD-TTA keeps Dice comparable while improving safety metrics, reducing HD95 from 70.96 mm to 64.55 mm (about 6.4 mm) and raising precision from 15.36% to 19.64% on MEN relative to the strongest baseline TCA.
  3. NVIDIA Research

    NVIDIA, Google DeepMind and EMBL-EBI openly released predicted protein-complex structures for more than 2,800 viruses, about 30% of them interactions never previously documented

    NVIDIA joined a coalition including Google DeepMind and EMBL-EBI to infer and openly release predicted 3D structures of protein complexes for more than 2,800 viruses using AlphaFold2 with the NVIDIA BioNeMo inference runtime, deposit them in the AlphaFold Database, and open-source the BioNeMo Structure Prediction Pipeline used to generate them, with about 30% of the protein interactions being entirely new to science.
  4. arXiv

    AGGRNet splits medical image features into informative and non-informative via a learnable threshold, reaching 5.48% higher accuracy than HiFuse-Small on Kvasir

    The work proposes AGGRNet, which embeds a Feature Extraction and Aggregation (FEA) module and a C2PCA block into a YOLOv11 classification backbone; spatial and channel attention plus a learnable threshold τ split feature maps into informative and non-informative parts that are then aggregated by cross-attention, yielding results above the compared SOTA models on five public datasets, including a 5.48% accuracy gain on Kvasir and a 2.2% gain on LIMUC.
  5. BMC Medical Informatics and Decision Making

    CleanSurvival uses reinforcement learning to auto-select preprocessing pipelines for survival analysis, improving predictive performance over simple baselines on real-world benchmarks

    The work presents CleanSurvival, a reinforcement-learning (Q-learning) based automated data-preprocessing framework tailored to time-to-event (survival analysis) models for censored data: it selects among combinations of data imputation, outlier detection, and feature extraction techniques to optimize performance for a Cox, random forest, neural network, or user-supplied time-to-event model; on real-world dataset benchmarks, this Q-learning-based preprocessing improved predictive performance relative to simple baselines, runtime behavior was condition-dependent and most clearly interpretable in the best-covered benchmark cells, and a simulation study showed effectiveness across different types and levels of missingness and noise.
  6. arXiv

    Replacing human labels with seven models across four abdominal CT datasets removed pretraining's dependence on label quality while direct deployment stayed quality-sensitive

    Across four abdominal CT datasets (WORD, AMOS, CT-1K, AbdomenAtlas), the authors generated pseudo-label variants with seven models (nnU-Net, MedSAM, TotalSegmentator, and four STU-Net sizes), then trained DynUNet under identical deterministic settings for in-domain training and for pretraining followed by fine-tuning; in-domain performance rose strongly and non-linearly with label quality and dataset volume partly compensated for quality, whereas fine-tuned models significantly outperformed a no-pretrain baseline in the vast majority of settings and pretraining label quality no longer clearly affected downstream results.
  7. arXiv

    Open-PMC-18M builds 18M medical image-text pairs via subfigure splitting and context summaries, lifting average retrieval by 27%

    Starting from the BIOMEDICA corpus, the authors used DAB-DETR-based subfigure detection, Qwen2.5-VL-32B-Instruct subcaption extraction, and Qwen2.5-14B-Instruct summarization of inline context to curate and release OPEN-PMC-18M, a dataset of 18 million subfigure-text pairs, then trained vision-language encoders evaluated on 6 retrieval and 19 zero-shot classification tasks across radiology, microscopy, and visible light photography, reporting an average recall of 21.64, a 27% relative improvement over the previous best model OPEN-PMC, and an overall zero-shot average F1 of 39.77.
  8. arXiv

    Retrieval-augmented anatomical guidance lets text-to-CT generation beat text-only baselines on fidelity, clinical consistency, and spatial controllability at once

    The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.
  9. Frontiers in Psychology

    CorticalG and PhysioG predict EEG and physiological responses to gravity, while Claude 3.5 Sonnet generates first-person narratives of altered-gravity awareness

    The study proposes a "gravity-awareness" framework in which two models trained on parabolic-flight literature — CorticalG, a Fourier-feature-augmented multilayer perceptron predicting g-dependent EEG band-power change, and PhysioG, eleven independent Gaussian processes predicting heart-rate variability, electrodermal activity and motor measures — are complemented by Claude 3.5 Sonnet generating first-person narratives of awareness at 0 g, lunar 0.16 g, Martian 0.38 g and hypergravity; results show alpha (DMN) and mu (SMN) suppression in microgravity, elevated beta (PFC) and gamma in hypergravity, a V-shaped electrodermal response (right-hand conductance rising over 200% at 1.8 g) and an inverted-U trunk-activity pattern (dropping over 50% at 0 g).
  10. arXiv

    Benchmarking seven video foundation models on 32,847 videos from 1,888 participants, VideoPrism ranks first on 10 of 16 tasks while V-JEPA2-SSv2 reaches 85.3% AUC on flip palm

    This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.

Page 20 · showing 10