Mira-Scene replaces sparse pose regression with pixel-aligned canonical coordinate maps, lifting 3D-IoU from 0.520 to 0.727 on BlendSwap
Synopsis
Mira-Scene is a compositional single-image 3D scene reconstruction framework that pairs a pixel-aligned Canonical Coordinate Map (CCM) with a scene-space Point Cloud Map (PCM) from monocular geometry estimation to form dense, bounded correspondences, recovers object transformations through robust geometric alignment, and jointly generates object geometry and CCMs with a multimodal diffusion transformer; across indoor, outdoor, synthetic, and in-the-wild scenes it improves layout accuracy over strong baselines, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D on BlendSwap while using limited open-source training data.
Interpretation
The paper replaces sparse, unbounded pose variables for object layout with dense, bounded correspondence recovery: the CCM maps each visible object pixel to a surface coordinate in the object's bounded canonical space, and when paired with a scene-space PCM from monocular geometry estimation, every valid pixel yields a canonical-to-scene correspondence, so object transformations are recovered by robust geometric alignment rather than neural pose regression. Prior compositional methods typically parameterize layout as sparse, unbounded pose variables such as translation, rotation, and scale, while holistic methods absorb placement into scene-level generation; the paper argues sparse poses are hard to regress under occlusion, perspective ambiguity, and long-tailed configurations, and that scene-level supervision is scarce. Because the CCM lives in bounded canonical space, it can be supervised directly from scalable object-level 3D data without scene-level layout annotations. The ablation (Tab. 4) compares three representations under the same geometry-layout co-generation architecture: Raw reaches 3D-IoU 0.365 and 2D-IoU 0.358, Coord Cube 0.379 and 0.381, and CCM with PCM 0.727 and 0.783; the paper explains that Coord Cube only densifies scene-space prediction and remains an unbounded target, so its gains are limited. Under matched training data, CCM with PCM still outperforms Raw and Coord Cube with 3D-IoU 0.537 and 2D-IoU 0.662.
The paper introduces a geometry-layout co-generation model: a Mixture-of-Transformers design in which a geometry expert generates object geometry in a 3D voxel latent space and a layout expert generates the CCM in pixel space, with the two modality-specific streams exchanging information through shared self-attention and with geometry and layout tokens embedded into a shared 3D positional space. The paper states it is the first to jointly denoise object geometry in a 3D latent space and pixel-aligned canonical coordinate maps in 2D image space within a Mixture-of-Transformers architecture; compared with simply concatenating tokens or fully separate models, this design preserves the distinct structures of the two modalities while exploiting the strong dependency between shape and layout. Architecture ablations show that removing joint attention drops 2D-IoU from 0.757 to 0.535 and raises CD from 0.017 to 0.070, while removing the shared positional embedding causes a smaller but consistent drop (2D-IoU 0.747, CD 0.019). The paper measures geometry-layout consistency by comparing the point cloud induced by predicted CCMs with the generated mesh, and notes that the strong Ours-(CCM+Mesh) result should be read as internal consistency between the two generated outputs rather than higher fidelity than the ground-truth mesh.
In full scene reconstruction evaluation, Mira-Scene's object geometry is competitive with the strongest baseline while layout accuracy improves markedly: on BlendSwap 3D-IoU rises from SAM3D's 0.520 to 0.727 and 2D-IoU from 0.672 to 0.783; on 3D-Future Scene 3D-IoU rises from 0.596 to 0.694. The paper emphasizes that SAM3D benefits from a much larger-scale data engine and substantially more object-level training data, whereas Mira-Scene uses only 60K open-source object assets; on object geometry Mira-Scene is slightly better on BlendSwap CD (0.021 vs. 0.027) and F-score (0.843 vs. 0.817), and slightly behind on 3D-Future Scene (CD 0.015 vs. 0.014, FS@0.1 0.845 vs. 0.866). Quantitative results come from Tab. 1 across the BlendSwap and 3D-Future Scene benchmarks, with metrics including CD, FS@0.1, EMD, 3D-IoU, ICP-Rot, 2D-IoU, and ADD-S; all methods receive the same scene RGB image and instance masks to factor out segmentation quality, and use the same normalization protocol and alignment procedure. The paper also reports that 3D-IoU falls from 0.752 for mostly visible objects to 0.635 for heavily occluded objects.
The paper contrasts its CCM with CUPID, another method that models 3D-2D correspondence: CUPID resembles forward rendering, storing for each 3D voxel center the pixel coordinate of its projection, whereas Mira-Scene resembles reverse rendering, directly recording a 3D coordinate for each 2D pixel, turning a one-to-many mapping into a one-to-one mapping and reducing correspondence ambiguity. The paper notes that CUPID can recover camera pose via a PnP solver and align a 3D object with a 2D image, but scene generation needs accurate 2D-pixel-to-3D-point correspondences to associate image observations with the point-cloud map; it also argues that predicting the 3D coordinates of a 2D token from 2D semantic features is easier than predicting the 2D coordinates of a 3D token from the same features. Under the CUPID correspondence protocol, CUPID-(Mesh+GT) reaches 2D-IoU 0.802, CD 0.047, FS@0.01 0.414, and FS@0.05 0.727, while Ours-(Mesh+GT) reaches 2D-IoU 0.795 but CD 0.023, FS@0.01 0.456, and FS@0.05 0.885. The paper notes CUPID's slightly higher 2D mask IoU while Mira-Scene's CD and F-scores are substantially better.
Perspective
The result targets compositional single-image 3D scene reconstruction: the input is one RGB image plus instance masks, and the output is a set of posed object assets, with object geometry generated in canonical space and layout recovered from dense CCM-PCM correspondences through robust geometric alignment. The paper reports applicability to indoor, outdoor, synthetic, and in-the-wild scenes, evaluated on the 3D-Future Scene and BlendSwap benchmarks, where BlendSwap covers indoor and outdoor environments plus realistic and cartoon-style appearances and serves as the primary benchmark for cross-domain generalization. The value of the method lies in freeing layout learning from scarce scene-level annotations so that object-level 3D assets can be used at scale for pre-training; pre-training uses 60K Objaverse objects and 1M object-centric views plus 20K background-completed photo-realistic object views, and fine-tuning uses 20K 3D-FRONT scene views. Downstream uses include removing, repositioning, reconfiguring, rigging, and animating individual objects in reconstructed scenes, and exporting object assets to interactive editing tools, embodied-AI simulators, and physics engines for perception, planning, manipulation, or physical simulation. When masks are not supplied, the paper uses a VLM-guided SAM3 segmentation pipeline to obtain visible instance masks, but all quantitative comparisons use ground-truth instance masks.
Appendix F lists several open questions: the current pipeline processes each object independently, so self-intersections can appear when objects are close or generated geometry is not accurate enough, as with the cat's tail intersecting the chair; the method requires an RGB image and corresponding object masks, and masks from real images are often noisy and can lead to floaters, so reducing mask dependence remains future work; the method mainly learns 2D-3D correspondence and relies on monocular geometry estimation for global scene geometry, which may produce thin, sheet-like point clouds for some 2D cartoon-style images; and current training data is substantially smaller than that of SAM3D and other models, leaving single-object generation quality relatively weaker. In addition, 3D-Future Scene evaluation selects only 50 examples because mask quality varies, and the SceneMaker comparison on BlendSwap follows a setting of at most 5 selected objects per scene, which bounds how far the conclusions extend. Quantitative results use ground-truth instance masks, so the practical effect of the automatic segmentation pipeline on final layout accuracy remains for readers to judge against their own settings.
