Single-image scene mesh generation with adaptive chunks and explicit 2D–3D conditioning outperforms all baselines on indoor and outdoor reconstruction
Synopsis
The work redesigns the object-centric 3D generator Trellis 2 with adaptive chunks that scale with camera distance, explicit 2D–3D conditioning that lifts image features onto observed surfaces while distinguishing free space from occluded regions, and supervision from about 4,000 synthesized outdoor scenes, generating a complete scene mesh from a single image and outperforming all baselines in geometric accuracy and perceptual quality on Tanks and Temples, ScanNet++, and 121 in-the-wild images.
Interpretation
An adaptive scene chunk representation is introduced in which nearby regions use smaller chunks to preserve detail while distant or tall structures use larger chunks to extend coverage, letting one generator trained on a canonical volume cover scenes of varying extent. Unlike fitting the whole scene into a single fixed-resolution volume or using overlapping chunks of one shared physical size as in GenRecon, the method increases chunk scale per depth row based on depth and local height. Ablations show adaptive chunking gives the best reconstruction scores on both indoor and outdoor scenes and runs three times faster than fixed 3 m chunks; on a large scene extending to about 200 m in depth and 100 m in height, fixed-size chunks would need roughly 200 chunks and exceed the memory of one NVIDIA A100 80GB, whereas the adaptive strategy covers it with four depth regions and 15 chunks in about 5 minutes.
Explicit 2D–3D conditioning lifts DINOv3 image features only onto the observed surface using a monocular point map, and encodes via a clipped signed-depth residual whether each 3D token lies in free space in front of the surface, on it, or in the unknown region behind it. Rather than letting every 3D token attend to the same image features through cross-attention, or tokenizing point maps and unprojecting image features directly, the method exploits the already-known token positions in adaptive chunking to condition each token directly. On both indoor and outdoor validation sets this conditioning outperforms SAM 3D-style and projection-only conditioning; qualitative inspection reports that the alternatives more often produce duplicated surfaces and miss geometry in occluded regions.
About 4,000 diverse outdoor scenes are constructed as additional training supervision, and training on these imperfect synthetic scenes is found to improve model performance. Existing scene datasets are largely indoor and procedural outdoor diversity is limited by hand-crafted rules; the work uses an agentic pipeline over 20 outdoor categories to generate reference images, object meshes, terrain, and layouts. With the conditioning mechanism fixed, adding outdoor data improves reconstruction on both indoor and outdoor scenes with larger gains outdoors; the trained model produces more detailed geometry and better preserves the input layout than the synthetic pipeline used to build its training data.
On Tanks and Temples, ScanNet++, and 121 in-the-wild images, the method outperforms all baselines in geometric accuracy and perceptual quality. Against seven baselines spanning iterative 3D completion, object composition, generated video followed by reconstruction, and direct scene reconstruction, the method leads on every metric. On Tanks and Temples it reduces Chamfer distance by 24% and improves F1 by 30% over the strongest baseline on each metric, and on in-the-wild images it is preferred in at least 73% of pairwise user-study comparisons.
Perspective
The method targets generating indoor and outdoor scene meshes from a single image, suited to settings such as content creation, virtual reality, and supplying explorable 3D worlds for physics simulators and robot learning; its chunk allocation relies on estimated point maps and camera intrinsics, and the training and inference pipeline builds on the public Trellis 2 and MoGe-3, so it applies to image inputs for which monocular geometry estimation is available. The authors state that future work will explore textured outputs.
The method generates geometry without textures and inherits Trellis 2's lack of watertight-mesh guarantees, occasionally producing holes that require additional mesh processing; results depend on estimated depth and camera intrinsics, whose errors can affect chunk placement, scene scale, and feature alignment; autoregressive generation may propagate errors from earlier chunks to later ones through reused overlap latents. In addition, the synthesized outdoor scenes are not perfect reconstructions, so how their shape and layout inaccuracies affect supervision quality remains an open question.
