Skip to main content
Back to timeline
arXivSource publication:

DeformSmith: Turning Text or a Single Image into Graspable Deformable Assets via a Physics Harness

Synopsis

DeformSmith presents a hierarchical agentic framework that, from text or a single image, progressively builds geometry, a physical model, material behavior, and robot interaction, using a shared physics-grounded harness to evaluate and revise candidates through simulation probes and robot pick-and-place feedback; across 39 cases its assets score higher on visual quality and physical plausibility than baselines including PhysGen3D, PhysGM, and PhysX-Omni, while also producing replayable manipulation data.

AI-generated editorial illustration: DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

Interpretation

It introduces an L0–L3 hierarchical construction pipeline: L0 reconstructs the mesh, Gaussian appearance, and a convex-decomposed collision proxy aligned in a common canonical frame; L1 initializes particle positions, volumes, masses, and contact conditions; L2 selects Young's modulus, Poisson's ratio, and damping for a neo-Hookean material via MPM simulation probes; L3 uses robot pick-and-place feedback to revise actions and materials. Prior single-image physical generation methods such as PhysGen3D and PhysGM mainly infer interactive physical representations from static input, whereas this work separates geometry, physical model, material, and robot interaction into layers with ordering dependencies, requiring later layers to revise while preserving quantities established earlier. The method is described in full, including the four layers, the contact velocity correction and friction formulation, and the reaction-force estimate; ablations show hierarchy versus a flat variant raises asset delivery from 83% to 93% and independent physics-test pass from 80% to 87% at comparable simulation cost (13 versus 13.7 calls).

It builds a shared physics-grounded harness in which LLM agents act as Planner, Designer, and Critic, a Prober runs geometric checks and simulation, and an Orchestrator applies rule checks against a contract, accepting or returning candidates within editable fields, parameter bounds, hard gates, and a revision budget. Simulation evidence is wired directly into the generation loop, so material parameters are not a one-shot inference but are observed by probes and repeatedly revised; in ablation, material-target satisfaction rises from 40% to 73% and hard failures fall from 17% to 7%, averaging 14 construction calls per asset. The ablation is assessed on independent deformation and recovery tests held out from candidate selection, with both variants starting from the same physical model and initial material proposal; the authors note the gains reflect both structured feedback and the additional simulation effort.

It brings robot manipulation into the generation loop: L3 plans approach, closure, lift, transport, release, and retreat, uses contact and deformation observations to diagnose action, material, and numerical issues, and records robot commands, particle states, contact observations, and task outcomes, retaining failed attempts as labeled diagnostic evidence. Basic physical probes do not cover behavior during grasping, transport, and release, so task-level feedback is added, and material changes must return to L2 for revalidation before robot confirmation under the same commanded plan, physical model, and force budget. The ablation uses six assets with five held-out conditions each; pick-and-place success rises from 40% to 67%, material pass rate from 67% to 83%, and joint success from 27% to 57%, with assets and actions frozen before held-out evaluation.

Across 39 cases (30 text-driven and 9 image-based from PhysGen3D's public project assets), it reports physical realism 0.70 and photorealism 0.58, each 0.24 above the strongest baseline PhysGen3D, with semantic consistency 0.82 versus 0.80; 40 participants rated anonymized, side-randomized pairs, with win rates of 68%–75% for physical plausibility, 88%–96% for visual quality, and 63%–76% for semantic consistency. Evaluation combines GPT-6 Astra automated ratings with blinded human preference, using matched input images and a shared interaction description: a five-second gravity drop from one quarter of object height, recorded at 30 fps from two views. Comparisons share the interaction description and physical duration, and automated ratings omit method labels; the authors state these results assess perceived quality under the tested interaction without establishing material-parameter accuracy.

Perspective

The work targets volumetric deformable solids; fluids and granular materials are out of scope. Results apply to the tested interaction setting, a five-second gravity drop from one quarter of object height plus simulated pick-and-place, and the authors state the evaluation concerns perceived quality rather than material-parameter accuracy. For a reader, this offers a path from text or a single image to simulation-ready, graspable assets with validation records and replayable interaction data; the authors propose using the generated result as an initialization combined with observation-driven methods such as DeformMaster and EMPM, using 3D point tracks from RGB-D video to constrain material-parameter refinement.

Material-parameter accuracy is not established, and the authors note physical parameters inferred from text or images require real-world validation; contact is modeled approximately, limiting coverage of complex deformable objects. Automated ratings and human preference both rest on the tested interaction and perceptual criteria, so whether conclusions hold under other tasks, object categories, or contact conditions remains to be seen. This reading covered the full text, but the specific visual content of Figures 2 through 5 can only be understood from the prose descriptions.

Sources