OTRetarget transfers contact relations from human demonstrations to a G1 humanoid, reaching 87% robot-object interaction Jaccard
Synopsis
OTRetarget jointly retargets robot and multi-object motion by representing surface interactions through signed distances, closest surface points and relative directions, transferring these quantities across human, robot and object geometries with entropic optimal transport, and enforcing them in a per-frame constrained inverse kinematics that optimizes robot and object poses jointly; on OMOMO it reaches a robot-object interaction Jaccard of 87% and a depth error of 8.7 mm versus 28% and 29.3 mm for OmniRetarget, and transfers to a physical Unitree G1 for two-handed box pick-and-place onto a table.
Interpretation
The method represents interactions as surface-level proximity triples (signed distance, witness point, relative direction) and uses them as interaction residuals inside a per-frame constrained inverse kinematics solved together with motion-style objectives. Prior skeleton-based retargeting describes surface interactions only partially, and interaction-mesh methods couple interaction accuracy and posture in a single metric; here interaction residuals and style objectives are expressed separately. The paper defines the residuals, probe sampling densities (1000 pts/m² on objects, 2000 pts/m² on human and robot) and weights, and compares against OmniRetarget, GMR and PHC on AMASS and OMOMO.
Object poses are treated as decision variables and optimized jointly with the robot, so robot-object, object-object and ground interactions are handled without rescaling the scene or the demonstration. Existing relational methods take the object pose as an input rather than a free variable, and OmniRetarget solves only for the robot and globally rescales the demonstration by the height ratio; here object trajectories adapt to the target embodiment. Table II shows that on OMOMO fixing the object slightly worsens style and ground contact, while freeing it recovers both with less than 1 cm drift; in the overhead-lift case a fixed object doubles style error and lifts the feet, whereas a variable object moves about 12.2 cm on average to a reachable height.
The optimal transport correspondence used for human-to-robot retargeting is extended to map interaction targets between the demonstrated object and a substitute object, so a single demonstration transfers to unseen object shapes and sizes. Earlier interaction transfer mostly targeted static partners or fixed objects; here the same transport framework maps interaction targets while keeping their metric meaning without rescaling. Table IV replaces the scanned box with eight unseen shapes, with recall and precision near the native box and depth error staying within about 12 mm; in the resizing experiment grip recall stays nearly constant while OmniRetarget's recall collapses on smaller boxes.
On OMOMO the approach reaches a robot-object interaction Jaccard of 87% and a depth error of 8.7 mm, versus 28% and 29.3 mm for OmniRetarget, and whole-body policies trained on the retargeted references transfer to a physical G1 for two-handed box pick-and-place onto a table. The paper reports interaction fidelity, runtime, object generalization and hardware transfer together, and indicates that the interaction residuals rather than the native scene or the variable object drive the gain. Table III reports medians over 4421 OMOMO sequences, and the scaled, fixed-object ablation retains most of the gain; the hardware part is a single-robot, single-clip demonstration that the authors state is not a hardware study.
Perspective
The result targets the retargeting stage that converts human demonstrations into humanoid whole-body loco-manipulation references, in settings with flat ground, a single robot (a 29-DoF Unitree G1 with fixed hands) and a single two-object capture; the formulation is kinematic and frame by frame, without explicitly modeling forces or optimizing over a time horizon, so the authors list more complex terrains and scenarios as future directions. For a reader, this means it can serve as an interaction-preserving tool at the retargeting stage and complements dynamics-based recovery methods.
The authors state that evaluation remains limited to one robot, flat ground and a single two-object capture, and that the formulation is kinematic, frame by frame, without explicitly modeling forces or optimizing over a time horizon; the hardware part is a single-robot, single-clip demonstration that the authors explicitly say is not a hardware study. In addition, some equations and table values appear as placeholders in the loaded text, and several specific numbers (such as runtimes and some error values) are not fully rendered in the body, so readers needing exact values should consult the original tables.
