Skip to main content
Back to timeline
arXivSource publication:

GOTT uses a reusable cross-embodiment contact-acquisition primitive to raise dexterous grasping to 78.7% and real-world task success from 35% to 80%

Related research and updates

Synopsis

GOTT introduces a reach-acquire-move framework in which a single cross-embodiment closed-loop contact-acquisition primitive turns a robot-agnostic object trajectory and reach specification into stable contact, after which a pose-conditioned controller tracks the desired object motion; in simulation the unified policy reaches 78.7% average grasping success versus 51.7% for CrossDex, transfers zero-shot to unseen Allegro variants at 72.1% versus 17.2% for an Allegro-only policy, and achieves 80.0% average task success on four real-world tasks versus 35.0% for open-loop baselines.

Source-provided article image: GOTT: Object-centric Dexterous Manipulation with a Reusable Cross-Embodiment Primitive
Figure 1 ·

Figure 1 : Contact-acquisition training and cross-embodiment policy. (A) The RL policy starts from randomized initial configurations and learns to establish stable contact using local hand-object observations. (B) Padding and binary masks map different hands into a shared observation and action space, allowing one policy to operate across hand embodiments.

arXiv

Interpretation

It proposes a reusable cross-embodiment contact-acquisition primitive that decouples where and how to establish stable contact from high-level intent. Prior methods either execute grasps open loop or learn contact acquisition jointly with trajectory execution, without exposing contact acquisition as a skill reusable across high-level sources and robot embodiments. Across 9,588 DexGraspNet assets and 54,149 object-scale-pose instances, the unified policy averages 78.7% versus 38.7% for BODex and 40.5% for DRO; real-world success is 90%, 90%, and 96% on three arm-hand platforms.

A canonical per-link representation lets one policy be shared across hand morphologies, and multi-morphology training drives zero-shot transfer. Compared with CrossDex's shared action space via human-hand retargeting, the per-link canonical representation raises average success from 51.7% to 78.7% under the same reward and PPO settings. On three unseen Allegro variants, the three-hand unified policy averages 72.1%, matching the per-variant oracle average of 72.2%, whereas an Allegro-only policy reaches 17.2%, indicating transfer comes from multi-morphology training.

Closed-loop contact acquisition substantially improves real-world end-to-end task success over open-loop grasp execution. Perception, VLM-based affordance reasoning, future-aware IK selection, and trajectory tracking are held fixed, isolating the contribution of contact acquisition. On four real-world tasks (trash cleanup, mustard pouring, whiteboard erasing, drill use) repeated 10 times each, GOTT averages 80.0% task success versus 35.0% for both BODex and DRO; DRO reaches 0/10 task success on trash cleanup because loose contact around the brush handle causes slipping.

Separating contact acquisition from workspace-specific reachability makes the same primitive more robust away from the training configuration. Embodiment-specific policies such as SimToolReal that jointly learn contact acquisition and tracking perform well near their training configuration but degrade sharply away from it. Over 600 simulated episodes, SimToolReal reaches 0.983 success near its training configuration versus 0.792 for GOTT, but below a height offset its success drops to 0.002 while GOTT achieves 0.914.

Perspective

The results concern tabletop manipulation of rigid objects, in settings where stable contact must be established before tracking an object trajectory, such as cleanup, pouring, wiping, and tool use. They let the high-level intent source (future-aware planning, external models, human demonstrations, generated videos, keypoints) be swapped while the contact-acquisition primitive and tracking backend stay unchanged, which is most valuable to system integrators who want to reuse one low-level skill across hands and arm platforms.

High-level planning and contact acquisition currently operate sequentially, so whether tighter feedback between them could enable online recovery from poor reach specifications remains an open question. Trajectory tracking relies on visual object pose estimation, which can fail under severe hand-object occlusion; incorporating tactile or visuotactile state estimation is a natural next step. The hand-object transform is treated as approximately fixed after grasping, and extending GOTT to handle changes in this transform would enable more precise in-hand manipulation such as unscrewing. In addition, real-world grasping experiments use 50 trials per platform and task experiments use 10 trials per task, so stability across more objects and scenes warrants further observation.

Sources