Public articles linked to the same research event.
arXiv The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.
The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.
The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.
The authors introduce PlaneCycle, a parameter-free, training-free, adapter-free 2D-to-3D lifting operator that lets a pretrained 2D backbone acquire 3D fusion by cyclically distributing spatial aggregation across the orthogonal HW, DW, and DH planes throughout network depth without modifying any pretrained parameter; using DINOv3 ViT-S/16, ViT-B/16, and ViT-L/16 on six 3D classification and three 3D segmentation benchmarks, PCg reaches 87.8 average AUC and 82.0 average ACC on ViT-B/16 under linear probing, surpassing slice-wise 2D and 3D-flattening baselines with paired t-test significance (p<0.05) on 5/6 datasets, and after full fine-tuning it matches standard 3D architectures, exceeding 3D flattening by up to 2.6 Dice points on segmentation while retaining 2D-level attention complexity.