GATOR reconstructs textured 3D objects and scene-relative pose from casual images with a generative and agentic framework
Synopsis
GATOR is a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more casual images: a local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, text-guided semantic conditioning supplies category names and object captions through stage-specific adapters for structure, geometry, and appearance generation, and the generated asset initializes a multimodal agent that refines structure and texture through an observation-guided edit-render-review loop; across synthetic objects, cluttered tabletops, and indoor scenes it achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations.
Fig. 1: Generative and agentic 3D object reconstruction from casual images. Left: Generative model inaccurately fills the strainer’s mesh bowl. GPT-6-Astra struggles with strainer ear position, and simplifies the handle and ear shapes. Our GATOR generates and refines the asset, preserving its shape and mesh. Right: GATOR reconstructs complete, textured, posed objects from real-world tabletop captures and room scans, which are composed using their predicted poses. Project page: https://research.nvidia.com/labs/lpr/gator/
arXivInterpretation
GATOR is introduced as a generative and agentic framework that reconstructs textured object assets and recovers their scene-relative pose from one or more casual images. It combines generative reconstruction with agentic refinement in a single pipeline aimed at inferring surfaces hidden by occlusions from sparse, uncertain observations, rather than only estimating shape from single or dense views. The summary reports strong geometric and appearance fidelity across synthetic objects, cluttered tabletops, and indoor scenes, with scene-relative pose recovered from sparse observations.
A local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning. This preserves scene context while distinguishing the target object from its surroundings, so reconstructions are scene-aligned rather than generated in isolation. The module is a central design element in the method description, and the summary ties it directly to preserving scene context and separating the target.
Text-guided semantic conditioning complements spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. Semantic cues supplement spatial cues and are adapted per generation stage instead of applying one conditioning scheme uniformly. The summary states that this conditioning complements the spatial cues and serves the structure, geometry, and appearance generation stages.
The generated asset initializes a multimodal agent that performs targeted structural and texture refinement through an observation-guided edit-render-review loop. The generated result serves as an instance-specific geometry and pose prior, so refinement is organized around the specific instance rather than generic post-processing. The summary describes the loop as observation-guided and states that it targets structural and texture refinement.
Perspective
The work targets reconstructing textured object assets and scene-relative pose from one or more casual images, in the settings described in the summary: synthetic objects, cluttered tabletops, and indoor scenes. Its value lies in supplying assets with instance-specific geometry and pose for downstream uses; the summary notes that scene-level simulation further demonstrates simulation readiness, making it especially relevant to simulation and content-creation pipelines.
The visible text is a summary and does not include specific evaluation metrics, dataset scale, baseline comparison details, or failure cases, so the quantitative level of fidelity and pose accuracy still needs confirmation from the full paper. In addition, the iteration cost and convergence behavior of the edit-render-review loop, and how text conditioning behaves when category names are unavailable, are open questions a careful reader may want to watch.
