Medicine & Health
466 items
TSegAgent achieves zero-shot tooth segmentation and FDI identification by pairing SAM3 with geometric reasoning, reaching mIoU 93.37 and TIR=1 of 87.17% on Teeth3DS while generalizing to a private dataset
The work proposes TSegAgent, which reformulates tooth instance segmentation and FDI labeling of intra-oral scanned 3D models as a zero-shot geometric reasoning problem: multi-view renderings with curvature heatmaps and the SAM3 text prompt "tooth" produce candidate masks, which are merged into face-level instance labels via IoU and containment relations, after which a geometry-aware vision-language agent performs non-tooth region identification, central incisor localization, full-arch classification, and error correction through multi-round conversation; it reports mIoU 93.37, TLA 96.40, TSA 96.76, TIR 97.46, and TIR=1 87.17 on Teeth3DS with 1200 3D tooth models, and mIoU 82.10, TLA 96.68, TSA 95.99, TIR 85.41, and TIR=1 51.
PromptGate lifts query purity in open-set federated active learning from about 60% to above 95% using federated learnable prompts
PromptGate introduces a client-adaptive vision-language gating module for open-set federated active learning (OS-FAL): it learns class-specific context (CSC) prompts on a frozen BiomedCLIP backbone, splitting them into global tokens aggregated via FedAvg and client-local tokens, uses VLM pseudo-labels to filter the unlabeled pool into a high-purity ID candidate pool before querying, and then hands that pool to any downstream active learning strategy; on the FedISIC and FedEMBED federated medical imaging benchmarks, static VLM prompting degrades to roughly 50% ID purity, whereas PromptGate maintains above 95% purity with 98% OOD recall.
In motor imagery BCIs, seven out-of-distribution detection methods failed due to intrinsic EEG variability, but Deep Ensembles and MC-Dropout reached up to 0.7 OOD detection ability for subjects with high on-task performance
This study used a Leave-One-Class-Out out-of-distribution detection setup in motor imagery BCIs, training a model on some classes and observing whether an unfamiliar movement class can be detected via increased uncertainty; it found that because of the high intrinsic variability of EEG signals, many users show higher uncertainty for familiar in-distribution classes than for out-of-distribution classes, so many OOD detection methods that perform well in other machine learning domains prove ineffective here, yet OOD detection performance correlates with on-task performance, and Deep Ensemble and MC-Dropout models achieved on-task AUROC above 0.9 and OOD detection ability up to about 0.7, showing that rejecting unfamiliar cognitive states becomes feasible when task performance is high.
BrainSTR models dynamic brain networks with spatio-temporal contrastive learning, validated on ASD, BD, and MDD with critical phases and subnetworks consistent with prior neuroimaging findings
The work proposes BrainSTR, a spatio-temporal contrastive learning framework that learns state-consistent phase boundaries via a data-driven Adaptive Phase Partition module, identifies diagnostically critical phases with attention, and extracts disease-related connectivity within each phase using an Incremental Graph Structure Generator regularized by binarization, temporal smoothness, and sparsity; a spatio-temporal supervised contrastive learning approach then leverages diagnosis-relevant spatio-temporal patterns to refine the similarity metric between samples and build a well-structured semantic space. Experiments on ASD, BD, and MDD validate its effectiveness, and the discovered critical phases and subnetworks provide interpretable evidence consistent with prior neuroimaging findings.
Spinverse inverts face permeabilities through a differentiable Bloch-Torrey simulator to reconstruct diverse microstructural interfaces on synthetic meshes
Spinverse presents a permeability-aware microstructure reconstruction method: on a fixed tetrahedral grid it treats each interior face permeability as a learnable parameter, optimizes face permeabilities by backpropagating a signal-matching loss through a fully differentiable Bloch-Torrey simulator, and recovers an interface by thresholding the learned permeability field; across a collection of synthetic voxel meshes it reconstructs diverse geometries and shows that sequence scheduling and regularization are critical to avoid outline-only solutions while improving boundary accuracy and structural validity.
A low-rank constraint turns the ultrasound video latent space into a readable cardiac-cycle trajectory, giving ED/ES frames without extra training
The work introduces LRM-Functa, which imposes a low-rank constraint on the time-resolved modulation vectors of VidFuncta (mt = v + Bβϕt), so that the latent space of cardiac ultrasound videos forms periodic spiral trajectories; this allows direct readout of end-diastolic (ED) and end-systolic (ES) frames without additional model training, keeps stable reconstruction and ejection-fraction prediction at the extremely low rank k = 2 on EchoNet-Dynamic, and generalizes to a POCUS cardiac view and to lung-ultrasound B-line classification.
PlaneCycle lifts DINOv3's 2D weights into a 3D model with no training and no adapters, reaching 87.8 average AUC under linear probing
The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.
TotalFM, an organ-separated 3D-CT foundation model, beats Merlin on 83% (25/30) of finding categories in zero-shot lesion classification while training at 32 batch/GPU
The study introduces TotalFM, an organ-separated 3D-CT radiology foundation model that uses TotalSegmentator and LLMs to automatically build roughly 340,000 organ-level volume-text pairs, then combines VideoMAE self-supervised pre-training with organ-wise contrastive learning; it reaches an average F1 of 0.708 in zero-shot organ-wise lesion classification (versus 0.515 for CT-CLIP and 0.650 for Merlin), a higher AUROC than Merlin in 83% (25/30) of finding categories, and report-generation performance comparable to Merlin while raising batch efficiency to 32 batch/GPU.
Robotic ultrasound drives real-time CBCT slice updating: USCorUNet cuts forward-backward residual by about 53% and updates a slice in 11.25 ms
The work proposes a deformation-aware CBCT updating framework that uses robotic ultrasound as a dynamic proxy to infer tissue motion: after hand-eye calibration initialization and LC2-based rigid refinement, a lightweight network, USCorUNet, estimates dense bidirectional deformation fields from adjacent ultrasound frames, which are spatially regularized and transferred to the CBCT reference slice, enabling real-time end-to-end CBCT slice updating without additional radiation exposure, validated on phantom and in vivo data for both deformation estimation and ultrasound-guided CBCT updating.
Page 18 · showing 10