Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ArticuTable reconstructs tabletop scenes with executable part-level articulation from a single RGB image, reaching the lowest LPIPS of 0.2398 and highest DINOv2 of 0.8991 on 150 reference images and ranking first in 85.59% of user-study evaluations

The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.