Public articles linked to the same research event.
arXiv The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.
The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.
The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.
The work presents ArticuTable, a single-image tabletop scene reconstruction framework in which GRAM converts imperfect monolithic proxy meshes into executable URDF articulated assets via MLLM-guided joint fitting and semantic state reasoning, while PSGSR progressively registers point-cloud, top-view, and input-view constraints with structure-aware semantic pose selection to resolve yaw ambiguity; on 150 generated reference images it achieves LPIPS 0.2398, DINOv2 0.8991, and CLIP 0.9265, outperforming ACDC, Gen3DSR, MIDI, and TabletopGen with the lowest collision rate of 0.3819%, ranks first in 85.59% of user-study evaluations, and releases 100 USDZ scenes directly loadable in Isaac Sim.