Dino U-Net, a frozen DINOv3 encoder with FAPM projection, reaches top segmentation across seven medical imaging datasets, with the 7B variant averaging 76.43% Dice
Synopsis
The work proposes Dino U-Net: a frozen DINOv3 foundation backbone as encoder, combined with a dual-branch DINO Adapter and a Fidelity-Aware Projection Module (FAPM), which outperforms seven baseline methods on seven public medical image datasets spanning endoscopy, ultrasound, microscopy, MRI, fundus and CMR modalities, and shows performance improving as the backbone scales from S to 7B.
Fig. 1. Architectural overview of the Dino U-Net. The encoder consists of a frozen DINOv3 backbone, a DINO Adapter, and a Fidelity-Aware Projection Module (FAPM), connected to a standard U-Net decoder.
· Page 4Interpretation
It proposes Dino U-Net, a hybrid architecture using a frozen DINOv3 backbone as encoder connected to a standard U-Net decoder, with a dual-branch DINO Adapter to bridge the domain gap between natural-image pretraining and medical images. Prior work integrating foundation encoders into U-Net frameworks largely built on the SAM family (e.g., SAM2-UNet); the authors describe the potential of DINOv3's dense features for medical segmentation as underexplored. The paper provides a full architectural description and diagram (Fig. 1) and compares against seven baselines on seven public datasets; the authors report statistical significance for most improvements (Wilcoxon signed-rank test, p<0.05).
It designs the Fidelity-Aware Projection Module (FAPM), which decouples features into a low-rank shared convolution branch and a specific convolution branch, uses the shared features to generate scaling factor alpha and shifting factor beta for affine modulation, and merges a refinement path with a shortcut path to preserve fine-grained detail during dimensionality reduction. The authors note that naive projections such as linear layers or 1x1 convolutions sacrifice fine-grained details, and FAPM is designed specifically for this reduction step. The ablation study (Table 4) shows that replacing FAPM with 1x1 convolutions degrades Dice by up to 0.79% and worsens HD95 by up to 1.75mm across scales; FAPM adds only 0.2-0.5M parameters to smaller models and reduces the parameter count by 0.8M for the 7B model.
Across seven public datasets, Dino U-Net variants outperform nnU-Net, SegResNet, UNet++, U-Mamba, U-KAN, Swin U-Mamba and SAM2-UNet overall; the 7B variant sets records on five of the seven datasets, with average Dice 76.43% (+1.87%) and HD95 18.76 (-3.04). The authors present this as the first systematic use of DINOv3 dense features for medical segmentation, reporting both regional accuracy and boundary accuracy metrics. Evaluation spans endoscopy, ultrasound, microscopy, MRI, fundus and CMR modalities, with dataset sizes ranging from 25 cases (MyoPS20) to 1044 cases (CellBinDB); Dice and HD95 are used, with significance testing reported.
The framework shows scalability: segmentation performance improves consistently as the backbone grows from S to 7B, and even the smallest S variant (5.11M parameters) exceeds the average performance of all baselines. The authors use this trend as evidence that medical segmentation benefits from the representational power of billion-parameter foundation models. Table 3 reports active parameter counts alongside average Dice/HD95 for each variant; the 7B variant has 228.97M parameters, and the authors flag its inference overhead as an open issue.
Perspective
The framework targets 2D medical image segmentation, applicable to organ, lesion and cell segmentation across endoscopy, ultrasound, microscopy, MRI, fundus and CMR modalities. For researchers and engineering teams wanting to reuse large-scale self-supervised representations at low training cost, the paper offers a reproducible architecture and public code (https://github.com/yifangao112/DinoUNet). The authors identify extending the framework to 3D volumetric segmentation as the next step.
The paper is set in 2D, so behavior on 3D volumetric data remains to be verified; the 7B variant leads on most datasets but carries markedly higher parameter count and inference overhead, and the authors list knowledge distillation as future work. Evaluation also relies on random 8:2 splits of public datasets, with no external cross-center or cross-device validation reported in the text. A reader of the abstract alone may not see how FAPM's benefit varies by model scale, which requires the ablation table to assess.
