Medicine & Health
466 items
HD-TTA chooses between competing 'compact or inflate' hypotheses, cutting HD95 by about 6.4 mm and raising precision by over 4% in cross-domain brain tumor segmentation
The work proposes Hypothesis-Driven Test-Time Adaptation (HD-TTA): with a frozen nnU-Net v2 backbone and optimization only over test-sample logits, a Gatekeeper first decides whether a case needs refinement, two competing geometric hypotheses (compact denoising vs. diffuse recovery) are generated in parallel, and a representation-guided selector picks the safest output using intrinsic texture consistency; trained on BraTS 2023 GLI and evaluated with strictly fixed hyperparameters on unseen pediatric (PED) and meningioma (MEN) target domains, HD-TTA keeps Dice comparable while improving safety metrics, reducing HD95 from 70.96 mm to 64.55 mm (about 6.4 mm) and raising precision from 15.36% to 19.64% on MEN relative to the strongest baseline TCA.
NVIDIA, Google DeepMind and EMBL-EBI openly released predicted protein-complex structures for more than 2,800 viruses, about 30% of them interactions never previously documented
NVIDIA joined a coalition including Google DeepMind and EMBL-EBI to infer and openly release predicted 3D structures of protein complexes for more than 2,800 viruses using AlphaFold2 with the NVIDIA BioNeMo inference runtime, deposit them in the AlphaFold Database, and open-source the BioNeMo Structure Prediction Pipeline used to generate them, with about 30% of the protein interactions being entirely new to science.
AGGRNet splits medical image features into informative and non-informative via a learnable threshold, reaching 5.48% higher accuracy than HiFuse-Small on Kvasir
The work proposes AGGRNet, which embeds a Feature Extraction and Aggregation (FEA) module and a C2PCA block into a YOLOv11 classification backbone; spatial and channel attention plus a learnable threshold τ split feature maps into informative and non-informative parts that are then aggregated by cross-attention, yielding results above the compared SOTA models on five public datasets, including a 5.48% accuracy gain on Kvasir and a 2.2% gain on LIMUC.
CleanSurvival uses reinforcement learning to auto-select preprocessing pipelines for survival analysis, improving predictive performance over simple baselines on real-world benchmarks
The work presents CleanSurvival, a reinforcement-learning (Q-learning) based automated data-preprocessing framework tailored to time-to-event (survival analysis) models for censored data: it selects among combinations of data imputation, outlier detection, and feature extraction techniques to optimize performance for a Cox, random forest, neural network, or user-supplied time-to-event model; on real-world dataset benchmarks, this Q-learning-based preprocessing improved predictive performance relative to simple baselines, runtime behavior was condition-dependent and most clearly interpretable in the best-covered benchmark cells, and a simulation study showed effectiveness across different types and levels of missingness and noise.
Replacing human labels with seven models across four abdominal CT datasets removed pretraining's dependence on label quality while direct deployment stayed quality-sensitive
Across four abdominal CT datasets (WORD, AMOS, CT-1K, AbdomenAtlas), the authors generated pseudo-label variants with seven models (nnU-Net, MedSAM, TotalSegmentator, and four STU-Net sizes), then trained DynUNet under identical deterministic settings for in-domain training and for pretraining followed by fine-tuning; in-domain performance rose strongly and non-linearly with label quality and dataset volume partly compensated for quality, whereas fine-tuned models significantly outperformed a no-pretrain baseline in the vast majority of settings and pretraining label quality no longer clearly affected downstream results.
Open-PMC-18M builds 18M medical image-text pairs via subfigure splitting and context summaries, lifting average retrieval by 27%
Starting from the BIOMEDICA corpus, the authors used DAB-DETR-based subfigure detection, Qwen2.5-VL-32B-Instruct subcaption extraction, and Qwen2.5-14B-Instruct summarization of inline context to curate and release OPEN-PMC-18M, a dataset of 18 million subfigure-text pairs, then trained vision-language encoders evaluated on 6 retrieval and 19 zero-shot classification tasks across radiology, microscopy, and visible light photography, reporting an average recall of 21.64, a 27% relative improvement over the previous best model OPEN-PMC, and an overall zero-shot average F1 of 39.77.
Retrieval-augmented anatomical guidance lets text-to-CT generation beat text-only baselines on fidelity, clinical consistency, and spatial controllability at once
The work proposes a retrieval-augmented text-to-CT generation method: given a radiology report, a pretrained 3D vision-language encoder retrieves the semantically nearest case from a reference corpus, and that case's anatomical annotation is used as a structural proxy injected into a text-conditioned latent diffusion model through a ControlNet branch; on the CT-RATE dataset (27,514 training and 1,818 test volumes), retrieval-augmented generation improves image fidelity (FID 3D 0.004), clinical consistency (CT-Net AUC 0.787), and spatial controllability (Dice 0.772, HD95 3.072) over text-only baselines, while additionally enabling explicit spatial controllability that such approaches inherently lack.
CorticalG and PhysioG predict EEG and physiological responses to gravity, while Claude 3.5 Sonnet generates first-person narratives of altered-gravity awareness
The study proposes a "gravity-awareness" framework in which two models trained on parabolic-flight literature — CorticalG, a Fourier-feature-augmented multilayer perceptron predicting g-dependent EEG band-power change, and PhysioG, eleven independent Gaussian processes predicting heart-rate variability, electrodermal activity and motor measures — are complemented by Claude 3.5 Sonnet generating first-person narratives of awareness at 0 g, lunar 0.16 g, Martian 0.38 g and hypergravity; results show alpha (DMN) and mu (SMN) suppression in microgravity, elevated beta (PFC) and gamma in hypergravity, a V-shaped electrodermal response (right-hand conductance rising over 200% at 1.8 g) and an inverted-U trunk-activity pattern (dropping over 50% at 0 g).
Benchmarking seven video foundation models on 32,847 videos from 1,888 participants, VideoPrism ranks first on 10 of 16 tasks while V-JEPA2-SSv2 reaches 85.3% AUC on flip palm
This study assembled a dataset of 32,847 webcam videos from 1,888 participants (727 with Parkinson's disease) across 16 standardized clinical tasks and systematically benchmarked seven video foundation models (VideoPrism, V-JEPA2 and its SSv2 variant, ViViT, VideoMAE, VideoMAEv2, and TimeSformer) using frozen embeddings with a linear classification head, finding that task saliency is highly model-dependent: VideoPrism excels in visual speech kinematics and facial expressivity, V-JEPA2 variants are superior for upper-limb motor tasks, and TimeSformer remains competitive for rhythmic tasks like finger tapping, with overall AUCs of 76.4–85.3%, accuracies of 71.5–80.6%, specificity up to 90.3%, and sensitivity of only 43.2–57.3%.
Page 20 · showing 10