Skip to main content
Back to timeline
arXivSource publication:

DPRD aligns ROI-masked displacement vectors so a student with ~5% of the teacher's parameters reaches 85.46% Dice on AMOS, edging out MedNeXt

Synopsis

The work proposes Displacement-Preserving Relational Distillation (DPRD), which in nnU-Net aligns batch-wise pairwise displacement vectors on ROI-masked case-level embeddings with relational scale normalization, and on ISLES 2022 and AMOS 2022 lets students using about 5% of the teacher's parameters and about 3% of its FLOPs outperform Logits KD, FitNet, RKD, and CIRKD, reaching 85.46% Dice and 11.72 mm HD95 on AMOS.

Source-provided article image: Displacement Preserving Relational Distillation for Robust Medical Segmentation
Fig. 1

Fig. 1. Overview of DPRD. RAFM obtains ROI-aware case-level embeddings, and DPRA aligns pairwise displacement relations between teacher and student representa- tions.

· Page 3

Interpretation

It introduces displacement-preserving relational distillation, using batch-wise displacement vectors with relational scale normalization to impose a scale-invariant local relational constraint instead of matching absolute activations. Compared with RKD, which collapses sample relations into scalar distances or angles, and CIRKD, which builds dense pixel-level relations for 2D images, DPRD models high-dimensional displacement direction and relative scale on compact case-level embeddings. The method is compared against Logits KD, FitNet, RKD, and CIRKD under the same nnU-Net training setting on ISLES 2022 and AMOS 2022, with ablations.

It develops a cross-architecture, ROI-aware, multi-stage distillation design that is memory efficient for 3D volumes and explicitly mitigates background-dominated supervision. ROI-Aware Feature Masking (RAFM) uses anatomical priors to focus alignment on task-relevant structures, and the mask is used only during training, with the teacher and projection layers removed at inference. Ablation shows that removing RAFM drops AMOS Dice from 85.46% to 81.68% and raises HD95 to 21.65 mm.

It consistently strengthens compact students on ISLES 2022 and AMOS 2022, with particularly pronounced gains on boundary-sensitive metrics. On ISLES, DPRD achieves the best Dice 75.16%, NSD 85.92%, and HD95 13.08 mm, with HD95 better than the MedNeXt teacher's 14.12 mm; on AMOS it reaches 85.46% Dice, 85.29% NSD, and 11.72 mm HD95, slightly exceeding the teacher's 85.37% Dice and reducing HD95 from 14.52 mm to 11.72 mm. Results follow the nnU-Net standard validation partitioning, with ISLES at 200 training/50 validation cases and AMOS at 240 training/60 validation cases, reported with means and standard deviations.

Ablations indicate that directional trajectory reconstruction and relative distance consistency provide complementary supervision. Removing trajectory mapping or distance regularization raises HD95 to 15.03 mm and 17.07 mm respectively, while the full framework achieves the best 11.72 mm. The ablation is run on AMOS 2022 with the MobileUNetV3 student, toggling RAFM, Traj., and Dist. components.

Perspective

The result targets settings that need to deploy 3D segmenters under constrained compute, such as localizing targets and at-risk structures in image-guided procedures; the method is built on nnU-Net, with a frozen MedNeXt teacher and MobileUNetV3 and PlainConvUNet students, and the teacher and projection layers are removed at inference so the deployed model matches the student in parameters and FLOPs. It applies to MRI stroke lesion and abdominal multi-organ segmentation tasks that have voxel-level annotations during training, and the authors note RAFM relies on ground-truth label masks, with future work exploring teacher-predicted ROIs or uncertainty-guided masks to relax this requirement.

Readers should still watch: ROI masking depends on ground-truth labels during training, so applicability in weakly-supervised or label-scarce settings remains unverified; evaluation is limited to the ISLES 2022 and AMOS 2022 benchmarks, and whether cross-center and cross-scanner appearance drift is fully covered awaits more data; the batch-wise displacement constraint is stochastic local supervision accumulated across sampled cases, augmented patches, encoder stages, iterations, and epochs, so its stability across batch sizes is worth attention; and although this is a full-text parse, some equations and table cells appear incomplete in the text, so reproducing the exact normalization and loss details would require the original paper and code.

Sources