Skip to main content
Back to timeline
arXivSource publication:

Imagine3D-LLM has multimodal LLMs imagine a compact 3D scene before answering, reaching 68.5 average on SPAR-Bench at 7B scale

Synopsis

The work introduces Imagine3D-LLM: a small set of learnable Gaussian summary tokens is appended after the image tokens, decoded into a compact 3D Gaussian Splatting representation, and trained jointly with a photometric reconstruction loss and standard next-token prediction, so that a multi-view multimodal LLM assembles a coarse 3D scene representation before answering; it consistently outperforms prior approaches across seven spatial reasoning and 3D scene understanding benchmarks, reaching 68.5 average on SPAR-Bench and surpassing the strongest spatially-aware baseline 3DThinker-7B by 5.2 points.

AI-generated editorial illustration: Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Interpretation

The model inserts a small set of learnable Gaussian summary tokens after the image tokens; hidden states at a middle LLM layer (layer 14) are decoded by a three-layer MLP into 3D Gaussian Splatting parameters, with each token decoding multiple Gaussians for 2592 Gaussians in total, supervised by rendering them back at the input viewpoints through differentiable rasterization with a photometric reconstruction loss. Prior routes to injecting 3D awareness either strengthened pixel-level cross-view correspondence or fused features from 3D geometry foundation models; here the MLLM itself assembles a compact, object-level 3D representation and conditions its answer on it. The paper provides a full method derivation and implementation details (built on LLaVA-Video-7B, 210 tokens per image, 32 input images, decoding at layer 14 of 28), and reports rendering quality of PSNR 17.61, SSIM 0.74, LPIPS 0.49 versus the ZipSplat teacher's 20.60/0.76/0.44, noting that reconstruction fidelity is not the design goal.

Because the number of summary tokens is far smaller than the number of image tokens, an information bottleneck forces content recurring across views to be merged into shared tokens; clustering the trained summary tokens with K-means and visualizing the Gaussians decoded from each cluster shows a single cluster reliably localizing one object (e.g., the chair), with no clustering or object-level supervision. This emergent object-level grouping turns the bottleneck design from an engineering trade-off into an observable mechanism, echoing the human spatial reasoning the paper cites: identify common objects across views, infer relative geometry between viewpoints, and assemble a coarse 3D layout. Evidence comes from clustering and attention visualizations and is qualitative; the paper also provides a quantitative probe: unsupervised semantic segmentation on the ScanNet scenes of the SQA3D evaluation set raises image-feature mIoU from a baseline of 26.54 to 35.43, approaching DINOv3's 37.62.

Reconstruction supervision is applied only at the summary tokens, yet the LLM's own image features become more 3D-aware: cross-view attention becomes sharper, and PCA visualizations show features organized semantically with the same object encoded consistently across views; following the 3DRS protocol, voxelizing ScanNet point clouds to form cross-view token pairs shows correspondence scores in the middle layers rising steadily during training. Cross-view correspondence was previously obtained through explicit supervision (such as 3DRS's correspondence distillation); here it emerges naturally from learning to reconstruct, without ever being directly optimized. The correspondence analysis is a quantitative probe (average cosine similarity), while attention and PCA are qualitative visualizations; the paper notes the improvement concentrates in the middle layers surrounding the Gaussian decoding layer.

Consistent gains across seven benchmarks: SQA3D EM 63.8, Real-3DQA EM 39.2, ScanQA CIDEr 109.3, Scan2Cap ROUGE 67.6, ScanRefer Acc@0.25 62.8, Multi3DRefer F1@0.25 60.2, and SPAR-Bench average 68.5 (low/medium/high 60.5/67.0/76.0), exceeding 3DThinker-7B by 5.2 points and surpassing the 72B-scale Qwen2.5-VL by more than 29 points. Against a controlled baseline sharing the identical backbone, training data, and schedule but omitting the summary tokens and the reconstruction and distillation objectives, every benchmark improves, indicating the gains come from the objective rather than extra data; ablations also show that feeding teacher tokens directly or supervising with distillation alone stays at baseline level. The controlled comparison covers six ScanNet-based benchmarks plus SPAR-Bench and includes ablations on decoding layer (7/14/21), number of summary tokens (1296/2592/5184), distillation, and training epochs; on cost, peak VRAM rises from 40GB to 44GB, and inference uses 20.43GB and 176.9ms, below VLM3R-7B which fuses external CUT3R features (25.29GB, 344ms).

Perspective

The result targets generalist multimodal models that take multi-view images as input and perform spatial question answering, situated reasoning, 3D dense captioning, and grounding; training and evaluation center on ScanNet-based indoor scenes and SPAR-Bench, with 32 frames per scene (the SPAR subset contains 2, 3, or 32 views). It enables follow-up work to embed compact 3D reconstruction as an inductive bias inside an MLLM without an external 3D model at inference; the teacher is used only during training and can be replaced by other token-level feed-forward 3DGS frameworks, and the Gaussian decoder head is invoked only when a reconstruction is explicitly requested.

Reconstruction fidelity is below the teacher model, and the paper explicitly states that high-fidelity rendering is not the goal, so the relationship between reconstruction quality and reasoning gains remains open. On training efficiency, convergence is slower without the teacher, and distillation mainly accelerates it. The representation budget is fixed; indoor scenes did not make it a constraint, but whether larger, more complex outdoor scenes need more Gaussians and dynamic token counts is still to be tested. In addition, the text underlying this summary includes the main paper, appendices, and ablation tables, but some visualizations (attention maps, PCA, clustered Gaussians, rendered views) appear as images and are known here only through their textual descriptions.

Sources