Skip to main content
Back to timeline
arXivSource publication:

OmniAct3D adapts perspective-trained vision foundation detectors to panoramic ERP, gaining 2.96 NDS over the previous best 3D detector on Spheriverse

Related research and updates

Synopsis

The work proposes OmniAct3D, a framework that adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic images by modeling spherical viewing rays and periodic spatial structure with an ERP-Ray Geometry Adapter, grounding each hypothesis in panoramic evidence and converting it into a structured geometric action via a Visual-Action Reasoning Chain, and re-encoding object regions at higher resolution for heading estimation with an Appearance-Guided Heading Expert; experiments report a 2.96 NDS improvement over the previous best 3D detector on Spheriverse, a 24.87 mAP improvement over the unadapted VFM baseline on PanoMMOcc, and 95-98% retention of same-configuration mAP by VARC under target-specific geometry adaptation.

Source-provided article image: OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection
Fig. 1 ·

Fig. 1: Comparison of monocular, multi-view, and ERP 3D detection paradigms. Monocular methods observe a limited field of view [ 3 ] , while multi-view methods rely on discrete perspective observations [ 4 ] . In contrast, our method enables coherent 360 ∘ 360^{\circ} perception from a single ERP image.

arXiv

Interpretation

OmniAct3D adapts perspective-trained Vision Foundation Model detectors to equirectangular projection (ERP) panoramic input while preserving transferable VFM priors. Existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; ERP encodes a continuous 360 scene in a single image, but organizes geometry and visual information differently, making direct transfer difficult. Abstract-level method description plus comparisons on Spheriverse and PanoMMOcc; the source is an abstract, so implementation details and ablation tables are not available.

The ERP-Ray Geometry Adapter (ERGA-Ray) mitigates ERP geometric mismatch by modeling spherical viewing rays and periodic spatial structure. Rather than reusing perspective-domain spatial inductive biases, it explicitly models spherical rays and periodic spatial structure to address the differing geometric organization of ERP. Module description in the abstract; network structure and parameters are not given in the abstract.

The Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence within scene-wide context and converts it into a structured geometric action. It shifts hypothesis verification from implicit feature regression to evidence localization plus structured geometric action, binding the reasoning process to panoramic evidence. Module description in the abstract, plus the reported result that VARC retains 95-98% of same-configuration mAP under target-specific geometry adaptation, suggesting reusable object-level 3D reasoning across sensing configurations.

The Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution to recover local cues lost under fixed token budgets for heading estimation. To counter local appearance cues lost under fixed token budgets, it separately re-encodes object regions at higher resolution to support heading estimation. Module description in the abstract; separate quantification of heading accuracy is not given in the abstract.

Perspective

The work targets panoramic 3D detection for mobile embodied agents, in settings where a single equirectangular image encodes a continuous 360 scene, with the goal of letting perspective-trained VFM detectors retain transferable priors on such input. Its value lies in offering an adaptation path for embodied perception systems that need localization and heading estimation under surround viewing, and in suggesting that object-level 3D reasoning may be reusable across sensing configurations. Applicability is bounded by the Spheriverse and PanoMMOcc experimental settings described in the abstract; behavior on other datasets, sensor configurations, and real deployments requires separate validation.

The abstract does not give per-module ablation contributions, training data scale, separate heading-estimation metrics, or the configuration range behind the 95-98% retention figure, so the relative importance of components is hard to judge. Behavior of the ERP geometry adaptation under extreme latitude distortion, the generalization boundary of VARC across more sensing configurations, and the token-budget versus resolution trade-off in AGHE remain open questions. Because this assessment is based on the abstract only, without the body figures and experimental details, the above judgments are limited to what the abstract states.

Sources