Skip to main content
Back to timeline
arXivSource publication:

Berkeley team's Morphometric Imitation retargets human hand-object interaction to three-, four-, and five-fingered robot hands, reaching 89.3% zero-shot success in 300 real-world trials

Synopsis

The work presents Morphometric Imitation, a three-stage framework that first kinematically retargets human hand-object interaction to robot hands of different morphologies while preserving demonstrated contacts via morphometric optimization (MMO), then uses residual reinforcement learning with object pose and contact information to produce dynamically feasible demonstrations, and finally distills them into visuomotor policies; across three robot hands and ten GRAB trajectories, MMO improves contact F1 over the strongest of five baselines by at least 8 points, downstream dynamic retargeting success by as much as 35 points, and the distilled policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects.

AI-generated editorial illustration: Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

Interpretation

MMO achieves the highest location-aware contact F1 and the lowest contact patch distance on all three robot hands through morphology matching followed by contact matching. Prior kinematic retargeting commonly relies on fingertip vector constraints or contact-proximity objectives; MMO instead scales the MANO hand model to the target robot hand's morphology and then re-optimizes the pose per frame to recover demonstrated contacts displaced by that morphological change. On Dex3, Allegro, and Sharpa with ten GRAB trajectories each, compared against five baselines (DexPilot, AnyTeleop, Position, Contact PyRoki, OmniRetarget), MMO leads the strongest baseline by at least 8 F1 points on every hand with lower patch distance; the paper also reports failure-adjusted scores and paired bootstrap intervals.

Better kinematic references carry through to downstream dynamic retargeting: with the residual RL formulation and hyperparameters held fixed, MMO references yield higher task success and lower contact-aware ADD. Earlier work rarely isolates the kinematic reference as the only variable under one shared dynamic retargeting formulation; this work does so. Each hand and object category is evaluated over 2048 simulated rollouts; MMO references give the best success rate and object tracking on all three hands, improving success over each hand's strongest baseline by as much as 35 points.

Object pose and contact information play complementary roles in residual RL: object pose alone deviates from the demonstrated grasp geometry, contact alone tracks the object trajectory less accurately, and combining both performs best. The work writes both object pose and contact information into observations, rewards, and termination conditions, and ablates each. Ablations show success drops on all three hands when contact information is removed and drops further when object pose is also removed; the paper attributes object pose to task-level motion guidance and contact to anchoring the demonstrated interaction.

The distilled visuomotor policies deploy zero-shot on hardware with 89.3% success across 300 trials and no failures from hard table collisions. The policies use no real-world training data, being distilled only from a privileged-state teacher in simulation, and still succeed on partially observable objects such as the transparent wineglass. On a KUKA iiwa14 with a Sharpa Wave hand and a RealSense L515, 10 categories with three physical instances each and 10 randomized initial poses per instance give 300 trials; failures stem mainly from pose-dependent localization error and insufficient pre-grasp opening rather than table collisions.

Perspective

The result targets tabletop manipulation by multi-fingered robot hands using reconstructed hand-object interaction as the demonstration source: MMO retargets at 180 frames per second for single-hand and 150 for bimanual trajectories on one RTX 4090, residual RL trains 60 to 90 minutes per policy, and hardware deployment uses a KUKA iiwa14 with a Sharpa Wave hand and a RealSense L515. It lets researchers transfer one human demonstration to a three-, four-, or five-fingered hand without collecting real robot data and obtain a directly deployable zero-shot visuomotor policy; for teams needing a unified cross-category policy or demonstrations taken straight from monocular video, the paper names both as next steps.

The paper's own scope notes that residual RL and visuomotor policies are currently trained separately per object category, hardware is validated on one arm-hand system, demonstrations come from ten motion capture trajectories rather than monocular-video reconstructions, and generalization to broader poses and shapes remains open. Real-world failures concentrate in pose-dependent localization error and insufficient pre-grasp opening, and the torus category falls well below simulation because its physical instances are thinner than the GRAB reference, indicating anisotropic scale randomization does not cover that geometric variation. In addition, contact evaluation is area-weighted over mesh vertices at a 5 mm tolerance, and the paper itself notes that a surface-resampling convergence study is needed before claiming physically calibrated millimeter accuracy; these are scope and open questions rather than refutations of the findings.

Sources