Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OpenBox aligns 2D vision-foundation instance cues with 3D point clouds to produce more accurate and cheaper 3D box annotations without self-training across three autonomous-driving datasets

OpenBox introduces a two-stage automatic annotation pipeline: in the first stage it uses a 2D vision foundation model and cross-modal instance alignment to associate instance-level cues from 2D images with the corresponding 3D point clouds; in the second stage it categorizes instances by rigidity and motion state and generates adaptive bounding boxes with class-specific size statistics, thereby producing high-quality 3D bounding box annotations without self-training and improving accuracy and efficiency over baselines on the Waymo Open Dataset, the Lyft Level 5 Perception dataset, and the nuScenes dataset.