SegDINO reshapes DINOv3 with lightweight scale modeling, reaching top segmentation accuracy across four datasets at 27.68M parameters and 51 FPS
Synopsis
The work presents SegDINO, which uses a frozen DINOv3-S encoder to collect intermediate features from layers 3, 6, 9, and 12, reorganizes same-resolution tokens into a pseudo multi-scale pyramid via Token Pyramid Adaptation (TPA), and applies Scale-Aware Decoding (SAD) for intra-scale refinement and top-down inter-scale propagation, alongside a new PanCT dataset of 284 pancreatic cancer patient CT scans; on PanCT and the public TN3K, Kvasir-SEG, and ISIC benchmarks, SegDINO outperforms U-Net, SegNet, R2U-Net, Attention U-Net, TransUNet, U-NeXt, and U-KAN in both DSC and HD95, with 27.68M total parameters and 51 FPS inference.
Fig. 1: An overview of the proposed SegDINO framework. A frozen DINOv3 en- coder extracts intermediate features from multiple layers. These features are first projected and reorganized by the Token Pyramid Adaptation (TPA) module to construct a pseudo multi-scale pyramid, introducing scale-aware representations. The resulting pyramid is then processed by the Scale-Aware Decoder (SAD), which performs intra-scale refinement and inter-scale information propagation using a residual refinement operator R. Finally, a lightweight prediction head produces the segmentation output.
· Page 3Interpretation
SegDINO reorganizes intermediate DINOv3 tokens into a pseudo multi-scale pyramid, replacing heavy decoders with scale modeling and achieving the best DSC and HD95 across four datasets. Prior work adapting DINO-family features for segmentation often stacks multiple multi-scale fusion modules or complex upsampling pipelines; this work instead applies 1x1 projection and strided-convolution resizing to frozen DINOv3-S features from layers 3, 6, 9, and 12 to build a pseudo pyramid at 1/4, 1/8, 1/16, and 1/32 resolutions. Compared with seven baselines on PanCT plus TN3K, Kvasir-SEG, and ISIC, Table 1 shows DSC 3.64% above the strongest baseline TransUNet on TN3K with HD95 lower by 7.50, and DSC gains of 4.64% and 2.41% over SegNet and U-KAN on Kvasir-SEG and ISIC.
Ablation shows TPA is the main contributor to performance gains while SAD adds lightweight refinement, with benefits varying by target scale. The Basic baseline projects multi-level encoder features to a unified channel space and concatenates them at a single resolution; M1 adds only TPA, M2 adds only SAD, and the full model combines both. In Table 2, PanCT goes from 0.7758 DSC for Basic to 0.8525 for M1 (about +7.7% DSC), 0.7906 for M2, and 0.8657 with HD95 2.61 for the full model; on TN3K differences are small (0.8318 to 0.8391).
The authors built the PanCT dataset for small-lesion segmentation, containing expert-annotated CT from 284 patients with confirmed pancreatic cancer. Collected from the Radiology Department, it uses 243 patients for training and 41 for internal testing; the training cohort has a median age of 59 years (94 female, 149 male) and the test cohort a median age of 56 years (20 female, 21 male), with independent annotation by two experienced radiologists and de-identified data under institutional ethical approval. Each 3D CT volume is converted to 2D axial slices for slice-level modeling and fair comparison with 2D baselines; the authors note the loss of inter-slice context and leave 2.5D/3D extension to future work, and the data are not publicly available but obtainable from the corresponding author on reasonable request.
SegDINO balances accuracy and efficiency, with competitive parameter count and inference speed. The model totals 27.68M parameters, 21.60M from the DINO-S backbone and 6.08M from the remaining components, which the authors state already outperforms most transformer-based counterparts. Inference reaches 51 FPS, surpassing most transformer architectures while remaining competitive with lightweight convolutional models; Fig. 3 presents the parameter, DSC, and FPS comparisons.
Perspective
The result targets medical image segmentation with 2D axial slices as input and a frozen DINOv3-S encoder, covering the evaluated tasks of thyroid nodules, colon polyps, skin lesions, and pancreatic tumors; for readers seeking multi-scale segmentation ability with fewer parameters and higher inference speed, TPA and SAD offer directly reusable modular designs, with code released at https://github.com/script-Yang/segdino_v2.
PanCT is not publicly available and only obtainable from the corresponding author, so external reproduction and cross-center validation remain to be seen; 3D volumes are converted to 2D slices, and the effect of inter-slice context is not yet evaluated; the authors list broader comparisons and ablations, generalization across more modalities and clinical scenarios, and 3D volumetric extension as future work, so results in those directions are not yet known; this is a full-text parse, but figures are rendered as text, so details of Fig. 2 and Fig. 3 cannot be fully verified from the text.
