Medicine & Health
464 items
IMaX swaps the marginal entropy term of mutual information for a Tsallis α-entropy, lifting accuracy by up to 7.3 points on ESCA and retinal datasets under long-tailed semi-supervised domain generalization
The work shows that semi-supervised domain generalization methods such as FBCSA and DGWM degrade substantially under long-tailed class distributions, and proposes IMaX: maximizing the mutual information between learned features and latent labels under supervision constraints from labeled samples, while replacing the standard marginal entropy term with a Tsallis α-entropy to relax the class-balance assumption; across two modalities (ESCA histopathology and diabetic retinopathy grading), three SSL frameworks and three SSDG methods, IMaX improves accuracy in all but one setting, with gains up to 7.3 points in the low-label regime (mL=5).
SAW conditions surgical video diffusion on four lightweight signals, cutting CD-FVD to 199.19 and lifting rare-action F1 from 20.93% to 43.14%
The work proposes Surgical Action World (SAW), which reformulates video-to-video diffusion as trajectory-conditioned surgical action synthesis conditioned only on four lightweight signals—a language prompt, a reference first frame, a tissue affordance mask, and 2D tool-tip trajectories—fine-tunes LTX-Video on a custom-curated dataset of 12,044 laparoscopic clips with a depth consistency loss, reaches CD-FVD 199.19 (vs. 546.82 for SurgSora) and FVD 224.28 on held-out test data, and demonstrates that augmenting rare actions with generated videos improves action recognition on real test data (clipping F1 20.93% to 43.14%; cutting 0.00% to 8.33%) plus initial feasibility of rendering tool-tissue interaction videos from simulator-derived trajectories.
StriMap integrates physicochemical, sequence-context, and interface-structural features to predict TCR–peptide–HLA recognition, and screening 13 million peptides from 43,241 bacterial proteins yielded candidate molecular mimics that activated T cells expressing an ankylosing spondylitis-associated TCR
The work presents StriMap, a unified framework that predicts TCR–peptide–HLA interactions by integrating physicochemical, sequence-context, and structural features at recognition interfaces, reporting state-of-the-art performance with improved generalizability; as a case study, the authors screened 13 million peptides from 43,241 bacterial proteins and identified candidate molecular mimics that were experimentally validated to activate T cells expressing an ankylosing spondylitis (AS)-associated TCR, with a top validated peptide enriched in patients with inflammatory bowel disease (IBD), suggesting potential shared microbial triggers.
BioGait-VLM reaches 68.1% accuracy on an 8-class gait benchmark and is chosen as the better model in 69.2% of blinded expert cases
The work proposes BioGait-VLM, a tri-modal (RGB vision, language, explicit biomechanics) framework built on a frozen InternVL3.5-1B, combining a Temporal Evidence Distillation branch with a Biomechanical Tokenization branch that projects 3D skeleton kinematics into language-aligned semantic tokens; on a subject-disjoint 8-class benchmark of 1,181 clips formed by merging public GAVD with a newly collected 30-patient degenerative cervical myelopathy (DCM) cohort, it reaches 68.1% accuracy and 52.9% macro F1 (98.1% F1 on DCM, 80.4% on Abnormal gait), and in a blinded review of 52 held-out clips four DCM-domain experts selected it as superior in 69.2% of cases (36/52, p<0.001).
Sparse autoencoders trained on 909,873 CT/MRI slices decompose medical foundation-model embeddings into language-describable concepts, recovering 87.8% of downstream performance with 10 features
Trained on 909,873 2D CT and MRI slices from the TotalSegmentator dataset, this study fits Matryoshka sparse autoencoders with BatchTopK sparsification on frozen embeddings from BiomedParse (biomedical) and DINOv3 (general-purpose) foundation models alongside a random-weight baseline, finding that sparse features reconstruct original embeddings with R2 up to 0.941, recover 87.8% of downstream performance with only 10 features (99.4% dimensionality reduction), preserve 97.7% of dense retrieval quality with five-feature fingerprints, correspond to monosemantic concepts expressible in language as verified by an independent LLM judge, and enable zero-shot language-driven image retrieval on a single clinical text query.
In simulations built on two randomized controlled trial datasets, Cox proportional hazards and random survival forest performance diverged by performance measure and proportional hazards assumption
The authors conducted a comprehensive neutral simulation comparison based on two reference datasets from randomized controlled trials, evaluating the Cox proportional hazards model and random survival forest for patient-specific survival probability prediction across multiple performance measures following TRIPOD recommendations, and found that conclusions based solely on the C index may not generalize to other aspects of predictive performance, that measures of overall performance may generally give more reasonable results, that the standard log-rank splitting rule for the random survival forest may be outperformed by alternative splitting rules particularly in nonproportional hazards settings, that the random survival forest performance suffered less in data with treatment-covariate inte
VA-Adapter lets an ultrasound foundation model guide the echo probe, cutting trained parameters about 33-fold while lowering guidance error below strong baselines
The work proposes a Vision-Action Adapter (VA-Adapter) inserted into the deep layers of a frozen ultrasound foundation model image encoder (EchoCLIP, USFM, BiomedCLIP) to online-inject understanding of individual 3D cardiac structure by encoding historical vision-action sequences; on a dataset of 178 adults, 356 expert scans and 1.31M image-action pairs, it reaches lower translation and rotation mean absolute error than strong probe-guidance baselines with roughly 2.61M-3.97M trainable parameters.
PIRTA pairs image-domain retrieval with text-domain augmentation, lifting ischemic-territory accuracy in 3D brain MRI reports by up to 57.2 points across cohorts
The study proposes PIRTA, a retrieval-augmented generation framework in which a 3D ViT encoder pretrained with large-scale MAE self-supervision retrieves clinically similar 3D DWI/ADC volumes in the image domain, and their paired clinician-authored reports ground LLaMA3-8B-Instruct report generation, avoiding explicit image-text alignment; on a 1,831-case multi-institutional internal cohort, a 580-case privacy-preserving external cohort, and the 206-case public ISLES benchmark, PIRTA achieves strong image-domain retrieval (internal mAP@1 of 94.0%) and consistently improves ischemic-territory accuracy, a clinically grounded surrogate for factuality, over direct image-to-text 2D baselines, leading the strongest 2D baseline by +57.2 / +34.5 / +30.6 multi-class accuracy points.
FetalAgents orchestrates specialized fetal-ultrasound models through multiple agents, achieving top performance across eight clinical tasks on external validation and auto-generating structured reports
The work proposes FetalAgents, a multi-agent system built on AutoGen in which a GPT-5-mini-driven Coordinator parses clinical intent and anatomical plane and dynamically dispatches expert agents that combine specialized vision models such as FetalCLIP, nnU-Net, USFM, SAMUS, and AoP-SAM via deterministic fusion rules, while a Summarizer consolidates outputs into structured reports; across eight clinical tasks (standard plane classification, brain plane classification, abdomen and stomach segmentation, AoP estimation, AC and HC measurement, GA prediction) on multi-center external datasets, FetalAgents achieves the best or most competitive results compared with specialized vision models, ultrasound foundation models, and general/medical MLLMs, and supports end-to-end video-stream keyframe ext
Page 16 · showing 10