Public articles linked to the same research event.
arXiv OpenBox introduces a two-stage automatic annotation pipeline: in the first stage it uses a 2D vision foundation model and cross-modal instance alignment to associate instance-level cues from 2D images with the corresponding 3D point clouds; in the second stage it categorizes instances by rigidity and motion state and generates adaptive bounding boxes with class-specific size statistics, thereby producing high-quality 3D bounding box annotations without self-training and improving accuracy and efficiency over baselines on the Waymo Open Dataset, the Lyft Level 5 Perception dataset, and the nuScenes dataset.
OpenBox introduces a two-stage automatic annotation pipeline: in the first stage it uses a 2D vision foundation model and cross-modal instance alignment to associate instance-level cues from 2D images with the corresponding 3D point clouds; in the second stage it categorizes instances by rigidity and motion state and generates adaptive bounding boxes with class-specific size statistics, thereby producing high-quality 3D bounding box annotations without self-training and improving accuracy and efficiency over baselines on the Waymo Open Dataset, the Lyft Level 5 Perception dataset, and the nuScenes dataset.
OpenBox introduces a two-stage automatic annotation pipeline: in the first stage it uses a 2D vision foundation model and cross-modal instance alignment to associate instance-level cues from 2D images with the corresponding 3D point clouds; in the second stage it categorizes instances by rigidity and motion state and generates adaptive bounding boxes with class-specific size statistics, thereby producing high-quality 3D bounding box annotations without self-training and improving accuracy and efficiency over baselines on the Waymo Open Dataset, the Lyft Level 5 Perception dataset, and the nuScenes dataset.
OpenBox introduces a two-stage automatic annotation pipeline: in the first stage it uses a 2D vision foundation model and cross-modal instance alignment to associate instance-level cues from 2D images with the corresponding 3D point clouds; in the second stage it categorizes instances by rigidity and motion state and generates adaptive bounding boxes with class-specific size statistics, thereby producing high-quality 3D bounding box annotations without self-training and improving accuracy and efficiency over baselines on the Waymo Open Dataset, the Lyft Level 5 Perception dataset, and the nuScenes dataset.