Frozen DINOv3 features injected into a 3D U-Net via Room-Lite mixing and calibrated fusion reach Dice 0.758 for 3DRA aneurysm segmentation and remove all cross-dataset failures
Synopsis
The work proposes DINO-3DRA, a dual-path framework that injects frozen 2D vision foundation model DINOv3-Small features into a trainable 3D U-Net backbone through Room-Lite spatial mixing and calibrated residual fusion, achieving state-of-the-art aneurysm segmentation on multi-centre 3D rotational angiography (3DRA) @neurIST data (233 patients, four institutions) with Dice 0.758, HD95 2.75 mm and a 13% gain over nnU-Net using only 5.72M trainable parameters; ablations attribute the gains to structured cross-dimensional transfer rather than loss design, and without fine-tuning on CADA (n=45) and SHINY-ICARUS (n=30) it reduces Dice<0.5 failures from 11.1% and 3.3% to 0%.
Fig. 1. Overview of DINO-3DRA. (a) Frozen 2D VFM (DINOv3-Small) features are fused with a trainable 3D U-Net backbone for vessel and aneurysm segmentation. (b) Multi-scale DINO features are refined by Room-Lite Mixers and fused into U-Net skip connections via Calibrated Fusion.
· Page 3Interpretation
A dual-path architecture in which frozen DINOv3-Small (ViT-S/16, 384-dim embeddings, pretrained on LVD-1689M via self-supervised distillation) processes all 64 axial slices independently and is combined with a trainable 3D U-Net backbone through calibrated fusion modules, adding roughly 120K parameters for 5.72M trainable parameters in total. Prior DINO adaptations for medical imaging were largely 2D, and naive slice-wise transfer failed to compete with optimised 3D architectures; this work brings 2D foundation-model semantics into 3D vascular segmentation, keeping pretrained knowledge in the frozen branch while the trainable branch learns volumetric context. Compared against U-Net++, Dual Attention, DeepVesselNet and nnU-Net on multi-centre @neurIST data (233 patients, four institutions, 4:1:1 split) with paired Wilcoxon tests at p<0.05; aneurysm Dice 0.758±0.234 versus 0.669±0.319 for nnU-Net.
A Room-Lite spatial mixer that restores the inter-slice continuity missing from slice-wise features using a learnable depth-positional bias B(z) and depthwise-separable 3D convolutions. DINOv3 processes slices independently, so the stacked volume lacks inter-slice continuity; this module rebuilds volumetric coherence in a parameter-efficient way, in contrast to DITR's setting which relies on RGB-D cameras and point clouds. In ablations, removing Room-Lite drops aneurysm Dice from 0.758 to 0.385±0.310 and raises HD95 from 2.75 mm to 22.08 mm (p<.001), the largest degradation among the variants.
Calibrated residual fusion that aligns the statistically different DINO and U-Net feature distributions via learnable scaling and a residual connection that preserves U-Net features when DINO features are unhelpful. Naive concatenation degrades performance (in preliminary experiments, +DINO naive lowered total Dice from 0.896 to 0.885); this design turns fusion from plain concatenation into a calibrated process with normalisation and residual injection. Removing calibration lowers aneurysm Dice to 0.492±0.301 and raises HD95 to 10.12 mm (p<.001); the random-feature control (ViT-S/16 with standard initialisation) reaches 0.628±0.279 and the DINOv2-small substitution 0.590±0.252, both significantly below the full model.
Cross-dataset evaluation shows that when trained on @neurIST and transferred without fine-tuning to CADA and SHINY-ICARUS, DINO-3DRA eliminates all Dice<0.5 failures (CADA 11.1% to 0%, SHINY-ICARUS 3.3% to 0%), with Dice +10.4% (p<0.001) and HD95 reduced by 75.8% (15.52 to 3.76 mm) on SHINY-ICARUS, while on CADA the mean gain is modest (p=0.607) but variance halves (sigma 0.186 to 0.093). Baseline architectures show catastrophic failures under heterogeneous protocols; this work frames stability, not only mean accuracy, as the main benefit of cross-centre transfer. The two external datasets comprise n=45 and n=30 randomly sampled cases and differ from @neurIST in scanner uniformity, acquisition protocols and label completeness; the authors state the conclusion still needs confirmation on larger external cohorts.
Perspective
The result targets intracranial aneurysm and cerebral vessel segmentation in 3D rotational angiography, with training and primary evaluation on multi-centre @neurIST data (233 patients, four institutions, resampled to 0.35 mm isotropic spacing, overlapping 64-cubed patches with stride 32) and external validation as no-fine-tuning transfer to CADA (n=45, aneurysm-only labels) and SHINY-ICARUS (n=30, vessel plus aneurysm labels). The method suits research and engineering settings that have expert-annotated 3DRA volumes and want stable aneurysm localisation under a small trainable-parameter budget; the authors note that homogeneous aneurysm confidence can offer a better geometric starting point for downstream surface meshing and hemodynamic modeling. The authors also state explicitly that robustness across protocols still needs confirmation on larger external cohorts.
The authors report that DINO-3DRA tends to under-segment large or morphologically atypical aneurysms, with ambiguous boundary voxels at the dome periphery where intensity gradients are subtle conservatively assigned to the vessel class, attributing this to training data dominated by compact saccular aneurysms while CADA includes fusiform morphologies, and to degraded boundary consistency for large aneurysms spanning multiple patches; the trade-off favours reliable localisation over volumetric completeness. The cross-dataset conclusions rest on randomly sampled cases from CADA (n=45) and SHINY-ICARUS (n=30), and the authors themselves call for confirmation on larger external cohorts. In addition, the aneurysm threshold of 0.005 was selected by grid search on the baseline and applied uniformly, so its optimality under different protocols, and the method's behaviour on non-3DRA modalities or other vascular beds, remain open questions to watch.
