Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

OmniAct3D adapts perspective-trained vision foundation detectors to panoramic ERP, gaining 2.96 NDS over the previous best 3D detector on Spheriverse

The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.