Skip to main content
Back to timeline
arXivSource publication:

DistScene treats the environment as an explicit component, generating decomposable 3D scenes from one image with a 32.2% relative scene-CD reduction indoors

Synopsis

DistScene is a framework for single-image compositional 3D scene generation that jointly generates the environment and multiple independent objects in a shared scene frame, refines each object in a local space, and distills pretrained object-generation priors into scene generation via automatically composed and rendered synthetic scenes; on indoor and outdoor benchmarks it improves scene-level spatial coherence over the evaluated baselines.

AI-generated editorial illustration: DistScene: Object-to-Scene Distillation for 3D Scene Generation

Interpretation

The work models the environment as an explicit scene component generated jointly with multiple independent objects in a shared coordinate frame, providing geometric context for object placement. Existing feed-forward methods largely represent a scene as a collection of object instances without explicitly modeling the surrounding environment, so object arrangements lack environmental geometric constraints; DistScene lets the environment latent act as a bridge anchoring object positions and coupling object latents. Ablation shows that adding the environment component raises MIDI-test object F-score from 29.98 to 40.98 and bounding-box IoU from 0.1798 to 0.2544, and on Gen3DSR-test reduces scene CD from 0.1367 to 0.1096 and raises F-score from 62.04 to 73.21.

Object-Centric Refinement refines each object in a normalized local space conditioned on scene context and places it back at its original scene location, recovering fine geometry without changing the global layout. Objects occupy only a small fraction of the global sparse voxel grid, limiting effective resolution; this stage refines each object independently in the same sparse voxel domain and uses a marker embedding to associate it with the corresponding region of the scene latents. Ablation shows refinement reduces MIDI-test object CD from 0.2037 to 0.1529, raises object F-score from 40.98 to 51.64 and IoU from 0.2544 to 0.3688; removing scene context degrades all metrics markedly.

Object-to-Scene Distillation uses a self-distilled data engine to automatically synthesize decomposable scene supervision, transferring pretrained object-generation priors to scene generation. Existing scene datasets are limited in scale and object diversity; the engine has an LLM produce a structured scene description, an object generator synthesize each object and the environment, and layout adjustment enforce contacts and avoid collisions, without any existing scene dataset or human annotation. Although both MIDI-test and Gen3DSR-test are out of domain for the synthetic data, the object-only architecture trained on it reduces Gen3DSR-test scene CD from 0.1936 to 0.1367 and raises F-score from 46.68 to 62.04.

On indoor and outdoor benchmarks, DistScene improves scene-level spatial coherence and object-level geometry over the evaluated baselines. On outdoor UrbanScene3D, CD-L1 drops from Extend3D's 0.0832 to 0.0772 and F-score rises from 0.680 to 0.722; on indoor MIDI-test, scene CD drops from 3D-Fixer's 0.1295 to 0.0877, and on Gen3DSR-test scene CD is 0.0958 with F-score 79.68%. Results come from quantitative comparisons on indoor MIDI-test, Gen3DSR-test and outdoor UrbanScene3D, plus an anonymized user study with 43 participants on 20 text-to-image-generated inputs where DistScene has the highest preference rate on all three criteria.

Perspective

The result targets single-image generation of decomposable 3D scenes, with training and evaluation covering indoor MIDI-test, Gen3DSR-test and outdoor UrbanScene3D, plus a user study on text-to-image-generated inputs. The method uses sparse voxels as a unified representation, with both scene generation and object refinement in that domain, so it directly supports component-level geometry output and object-level editing; the 512-resolution plus refinement configuration offers an efficiency-quality trade-off at 5.7 s scene generation, 3.3 s per refined object and 18.0 GB peak memory, suited to speed-sensitive settings. For downstream tasks needing explicit environmental context, decomposable components and local detail, such as simulation and robotics-related scene construction, this representation provides a reusable basis.

Indoor environment assets in training generally omit the ceiling, so some reconstructions may show open-top rooms; outdoor environments span large extents, spreading limited scene-frame voxels over a wide region and potentially losing fine detail or containing holes, and the outdoor training set is smaller than the indoor one. These fall within the scope of the current data-generation and representation setup, pointing to future attention on enclosed indoor environment assets and large-scale scene representations. In addition, this material is a full-text parse; if figure and table details are not fully rendered, the exact presentation of qualitative examples and the user study can still be checked further.

Sources