Public articles linked to the same research event.
arXiv The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.
The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.
The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.
The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.