ArticuTable reconstructs tabletop scenes with executable part-level articulation from a single RGB image, reaching the lowest LPIPS of 0.2398 and highest DINOv2 of 0.8991 on 150 reference images and ranking first in 85.59% of user-study evaluations
Related research and updatesSynopsis
The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.
Figure 1: Single-image reconstruction of articulated tabletop scenes. Given one RGB image (left), ArticuTable reconstructs an input-view-consistent 3D scene containing rigid and articulated objects. Renderings of the same scene under different joint configurations (right) demonstrate part-level articulation while preserving the arrangement of object instances.
arXivInterpretation
GRAM recovers executable articulated assets from imperfect image-generated monolithic meshes without articulation-specific training or fine-tuning. Existing articulation modeling methods are largely trained and quantitatively evaluated on curated articulated-object datasets with well-defined part structures and articulation annotations, whereas image-generated meshes often have fused part boundaries, distorted contact geometry, and incomplete internal structures; GRAM replaces reliance on local geometry alone with MLLM-guided joint fitting and semantic state reasoning. Evaluated on all 243 Lightwheel assets under the official protocol, GRAM reaches part-matching precision 91.15, joint angular error 17.26 degrees, and axis-location error 0.034, the best among the evaluated mesh-conditioned methods; ablations show MLLM-guided joint fitting reduces angular error by 57.1% and axis-location error by 20.9%.
PSGSR recovers an input-view-consistent scene layout through three-stage progressive registration over point-cloud, top-view, and input-view constraints plus structure-aware semantic pose selection. Existing methods commonly optimize object poses and scales directly against the input view, yielding a non-convex objective that depends strongly on initialization; PSGSR initializes each stage from the preceding estimate and uses structure-weighted bidirectional semantic correspondences to pick yaw hypotheses for differentiable refinement. Ablations show the full model achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, better than point-cloud alignment only (0.3439/0.8102/0.9077) and the variant without SASPS (0.2543/0.8762/0.9205); the top-view stage uses inexpensive 2D rotations to restrict input-view search to residual symmetric hypotheses, with the full pipeline at 179.4 seconds versus 374.8 seconds for batched front-only SASPS and 595.3 seconds with sequential rendering.
On 150 generated tabletop reference images, ArticuTable leads four baselines across input-view consistency, GPT-5 evaluation, and physical validity. Relative to the strongest baseline TabletopGen, LPIPS drops from 0.3935 to 0.2398, DINOv2 rises from 0.7788 to 0.8991, the GPT-5 average rises from 5.378 to 6.240, the object-intersection rate falls from 1.4488% to 0.3819%, and the fraction of scenes containing at least one inter-object intersection falls from 16.00% to 10.67%. All baselines share inputs and views and, where applicable, use the same TRELLIS.2 and GPT-5 as ArticuTable; physical validity is evaluated at the initial scene configuration, excluding pairs involving the supporting table.
The authors release ArticuTable-100, a collection of 100 simulation-ready tabletop scenes directly loadable in Isaac Sim, containing visual meshes, collision geometry, and executable articulated assets. The collection is distributed as self-contained USDZ packages, and articulated links use per-link SDF colliders to preserve interaction-relevant concavities such as open containers, internal compartments, and through-holes rather than a single convex hull. Package-integrity and initial-state checks cover all 100 scenes; Isaac Sim demonstrations include object removal, insertion, and rearrangement, plus gravity settling, toolbox transport of unattached contents, and a lid joint that remains operable after relocation.
Perspective
The result targets the setting of reconstructing tabletop scenes from a single monocular RGB image, applicable when objects can be instance-segmented, the table can be detected, and a synthesized top-view hypothesis is available; outputs are URDF and self-contained USDZ packages directly loadable in Isaac Sim for object removal, insertion, rearrangement, and physical interaction. For downstream embodied-manipulation research, it provides a 100-scene collection with visual meshes, collision geometry, and executable articulated assets, plus a pipeline that converts generated meshes into executable assets without articulation-specific training. The method relies on fixed pretrained components including GPT-5, SAM 3, TRELLIS.2, P3-SAM, VGGT, and a frozen DINOv2, with no task-specific fine-tuning.
The reconstructed scene is rendered with an estimated pinhole camera whose projection model and parameters may not fully match those of the input image, and under strong off-axis viewing such discrepancies can become more visible near image boundaries; Fig. 7 shows the reconstructed cup exhibiting less projected tilt than the input cup. Scene-level metrics such as LPIPS, DINOv2, and CLIP primarily measure overall image similarity, so local orientation discrepancies may have limited effects on their scores and an incorrect orientation can sometimes score higher, meaning these metrics are best read alongside qualitative comparisons. In the manual audit, 82.8% of the 418 judgeable assets pass, with the remainder partial or failing; failure causes include motion-range errors, part-segmentation errors, implausible axes or pivots, topology or joint-type errors, collision or detachment, and missing articulation, and the two evaluators reach 83.5% exact agreement at the asset level with quadratically weighted kappa of 0.474 for structure and 0.659 for motion. These are scope and open questions: joint refinement of camera parameters and object transformations, explicit lens-distortion modeling, and behavior on a broader range of real images remain for future work to examine.
