Skip to main content
Back to timeline
arXivSource publication:

PlaneCycle lifts DINOv3's 2D weights into a 3D model with no training and no adapters, reaching 87.8 average AUC under linear probing

Synopsis

The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.

AI-generated editorial illustration: $$PlaneCycle$$: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

Interpretation

PlaneCycle is a parameter-free, training-free, adapter-free operator that gives a pretrained 2D backbone intrinsic 3D fusion capability by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth. Prior 2D-to-3D lifting required retraining, adapters, or architectural redesign, or, as with ACS convolution, was restricted to CNN architectures; PlaneCycle is presented as architecture-agnostic (compatible with both CNN and ViT), adds no new parameters, and leaves pretrained weights untouched. The method is specified in Algorithm 1 with a complexity analysis: per-layer self-attention cost matches slice-wise 2D, a D-fold reduction versus full 3D flattening; the paper reports 16.3 h versus 36.2 h training time on ViT-L/16 on LIDC.

With no training at all, the lifted models already produce usable 3D representations: frozen-backbone features align well across the HW, DW, and DH planes and score higher FeatDice than both 2D and 3D baselines. The paper reports that slice-wise 2D inference is well-structured within the chosen plane but inconsistent across slices, while naively converted 3D models show weak, poorly aligned representations on all three orthogonal planes before retraining; PlaneCycle yields cross-plane-consistent 3D features without additional supervision. Based on PCA visualizations (Fig. 1) and the newly defined FeatDice metric (Table 3, zero-training features), e.g. on ViT-S/16 PCg reaches 42.6 on LIDC versus 36.9 for 2D and 13.9 for 3D.

Under linear probing, PlaneCycle achieves the best result on nearly all settings across six 3D classification and three 3D segmentation benchmarks, outperforming slice-wise 2D baselines and strong 3D counterparts and approaching fully trained models. The paper reports that on ViT-B/16, PCg surpasses R-ACS by 3.0 AUC and nearly 6.0 ACC on average; in segmentation linear probing, PCg averages 64.8 Dice on ViT-S/16, 65.2 on ViT-B/16, and 65.6 on ViT-L/16, all above 2D and 3D flattening. Results are averaged over five runs, and paired t-tests (Fig. 3) confirm significance at p<0.05 on 5/6 datasets under linear probing.

Under full fine-tuning, PlaneCycle matches standard 3D architectures while retaining 2D computational efficiency. The paper notes that 3D flattening relies heavily on end-to-end fine-tuning to realize its potential, whereas PlaneCycle remains competitive or better after fine-tuning: up to 2.6 Dice points above 3D flattening in segmentation, and on classification ViT-B/16 PCg averages 91.2 AUC versus 91.1 for 3D flattening, exceeding the Transformer-based volumetric model ViViT by up to 2.6 AUC. Full fine-tuning results in Tables 2 and 3 are averaged over five runs, with training-cost comparisons given (16.3 h versus 36.2 h on ViT-L/16 on LIDC).

Perspective

This work targets researchers and engineering teams who need to apply existing 2D foundation models to volumetric data, especially in medical imaging; it is validated on six MedMNIST+ 3D classification datasets (Organ, Nodule, Fracture, Adrenal, Vessel, Synapse, spanning CT, MRI, and electron microscopy) and on the LIDC and MMWHS segmentation benchmarks, with DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 backbones, all on a single NVIDIA H200. It is positioned as a complementary operator to 3D pretraining and adapters rather than a replacement, so it fits best where 2D weights already exist, 3D fusion is wanted at very low extra cost, and supervised fine-tuning remains available.

The paper states that systematic comparisons between CNN and ViT backbones have not yet been conducted, and that combinations with 3D pretraining and adapters such as LoRA remain to be explored; large-scale potential is not yet validated, with only preliminary results indicating that lifting DINOv3-7B is feasible. Segmentation experiments intentionally use a shared lightweight decoder to isolate the lifting effect, and the authors note the simple decoder and the 16x downsampling in ViT may limit spatial recovery. In addition, FeatDice is newly defined to assess volumetric feature coherence and has no established standard for direct comparison, and the plane schedule uses a four-operator HW to DW to DH to HW cycle, leaving other schedules an open question.

Sources