Skip to main content
Back to timeline
arXivSource publication:

USDCraft lets a pretrained LLM write executable programs grounded in partial geometry to rebuild articulated 3D assets, holding sim-to-real loss to at most 10 points in real robot manipulation

Synopsis

USDCraft formulates articulated 3D asset reconstruction as programmatic modeling grounded in partial geometric evidence: a pretrained LLM without task-specific training writes, compiles, and revises executable asset programs from a source mesh and reference image, where source geometry analysis converts the mesh into a metric text description separating observed surface from unknown space, iterative geometric rechecking compares each candidate against the source in that same representation to point to program edits, and visual feedback plus physical authoring guidance complete the process, yielding articulated USD assets with explicit physical properties that load into Isaac Sim without manual adjustment.

Source-provided article image: USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation
Figure 1 ·

Figure 1: USDCraft enables generation, reconstruction, and real-to-sim-to-real manipulation.

arXiv

Interpretation

It formulates articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence, solved by pretrained LLMs without task-specific training; parts, joints, and physical properties are explicit program parameters, so a discrepancy found during inspection can be fixed by a local edit, and the program can introduce parts and connections the source never observed. Prior mesh-based methods can only partition the given surface, so fused parts stay merged and missing parts cannot be added; image-only programmatic generation lacks measurements of the target instance, so dimensions and part placement may not match. USDCraft uses measured geometry for both construction and revision while retaining the freedom to rebuild missing geometry and separate fused parts. On USDCraft-bench, USDCraft–Sol and USDCraft–Astra outperform all four baselines (ArtLLM, SIMART, Particulate, Articraft) on every metric; USDCraft–Astra results are means over three runs whose variation is far smaller than these margins.

Source geometry analysis converts the source mesh into a compact metric text representation: surface support is sampled on a regular grid in a shared metric frame and serialized as axial slice maps along the vertical axis with local metric summaries of bounds, protrusions, openings, and observed interior fittings, while every cell outside the sampled support is marked unknown rather than empty or solid. Rendered views discard metric scale and hide interior structure, and raw vertex or point lists exceed a practical context budget; distinguishing unknown from empty or solid keeps a hole in a scan from implying free space and the inside of a closed shell from implying solid material, so the agent can complete such regions from the image and the object's function. Ablations show source geometry analysis brings the next largest gain after the authoring harness, and removing it causes the largest drop in part recovery, geometric overlap, and joint accuracy, since later rechecking can correct a candidate but cannot supply the measurements needed to plan it.

Iterative geometric rechecking re-encodes each candidate on the source grid and returns source coverage plus two discrepancy sets: observed source surface the candidate misses serves as a correction signal, while differences where the source is unknown remain open questions resolved with the image and the object's function; first-surface maps from six axis-aligned directions and selected-region depth comparisons further decompose discrepancy into a median offset and the 90th percentile of the residual after removing it, indicating front–back shift versus shape difference. Coverage alone cannot tell whether a part is misplaced or wrongly shaped, and the two cases call for different edits; this decomposition lets the agent trace each discrepancy back to the part of the program that produced it and recompile and recheck the affected geometry after each revision. In the paper's example, a handle shows a median offset of only 0.64 mm against a shape difference of 3.41 mm, indicating its profile rather than its placement needs revision; ablations show iterative rechecking mainly improves geometric fidelity and joint placement.

In real-to-sim-to-real manipulation, policies trained on USDCraft assets lose at most 10 points from simulation to reality on drawer opening, toaster-lever pressing, and toaster switching, whereas image-only Articraft loses 60–80 points and Mini Workflow varies widely. A policy learns where to grasp and push from the asset, so a handle, lever, or switch that is misplaced or missized in simulation sends the real robot to the wrong location; structural completeness must come before geometric accuracy, as Particulate's incomplete reconstructions support only lever pressing. For each method–task setting whose asset supports the required interaction, 200 simulation demonstrations are collected and a Diffusion Policy is trained for 80 epochs with batch size 128, evaluated in 20 simulation and 20 real-world trials with randomized object positions and initial robot configurations.

Perspective

The work targets robot learning and real-to-sim-to-real transfer that need simulation-ready articulated assets: inputs are a source mesh, a reference image, and an optional text request, and the output is an executable program compiling into an articulated asset with rigid parts, joints defined by type, axis, origin, and limits, and physical parameters such as masses and contact friction, loading into Isaac Sim without manual adjustment. The same authoring toolkit supports text- or image-conditioned generation without a source mesh, producing USDCraft-10k spanning more than 500 everyday object categories. Evaluation covers 60 agent-generated assets plus a 40-object extension in USDCraft-bench, 243 human-modeled articulated objects in Lightwheel, and real manipulation scenes with a drawer, toaster, and box. The +G setting shows users can steer part and joint definitions through plain text without retraining or changing the method.

Fine lattice structures such as racket strings and shopping-cart wire frames remain hard to reconstruct faithfully, since their closely spaced members require accurate local geometry and consistent connectivity; flexible members can currently be annotated and given rest geometry, but how they bend, stretch, or respond to contact has not been validated in Isaac Sim, and canopy tension and rigid–flexible coupling are not modeled. Generation quality is scored by a model-based judge, some configurations use a single run, and run-to-run consistency is reported only for the Astra configuration over three repeats. Benchmark references are annotated by manually merging P3-SAM initial segmentations and annotating joints; source-exclusion comparisons show the advantage does not depend on USDCraft-generated or agent-generated references, but the influence of the annotation procedure itself remains an open question. Absolute scale is retained in the reference-frame evaluation, where image-conditioned methods drop noticeably, suggesting the effect of scale assumptions on deployment deserves attention.

Sources