Skip to main content
Back to timeline
arXivSource publication:

PatternDex guides bimanual dexterous manipulation with object-dictated interaction patterns, raising Allegro success from 52.2% to 92.8%

Related research and updates

Synopsis

PatternDex learns an embodiment-agnostic "interaction pattern" token sequence from human-object demonstrations and, combined with a target robot hand description, decodes wrist motions and contact points that fit that hand to guide a reinforcement learning policy; on 6 unseen demonstrations from the ARCTIC dataset it reaches a 92.8% mean success rate with Allegro hands versus 52.2% for the baseline ObjDex, transfers to Shadow, Paxini, and XHand hands with 84.0%, 73.5%, and 71.6% after fine-tuning alone with a frozen pattern encoder, and opens a real microwave in 7 of 10 trials.

Source-provided article image: PatternDex: Learning Interaction Patterns to Guide Reinforcement Learning of Bimanual Dexterous Manipulation of Articulated Objects
Fig. 4 ·

Fig. 4 : Qualitative results of the guided policy. (a) Guided by the estimated wrist motions and contact points shown in the top row, Allegro hands open the microwave, the laptop, and the notebook. (b) Shadow, Paxini, and XHand hands succeed on the same tasks with the interaction pattern transferred from the Allegro hand.

arXiv

Interpretation

It introduces the "interaction pattern," a latent representation that encodes the relationship between a desired object trajectory and the wrist motions achieving it as a token sequence, where each token encodes the relative translation and rotation of the wrist with respect to the object at a frame. Prior hand-object motion work such as OMOMO and CAMS generates human hand motion and has not been applied to robot hands of a different embodiment; PatternDex represents this correlation explicitly as an embodiment-independent token sequence. The encoder combines a spatial MLP with a temporal network of self-attention heads and is trained on 229 human-object demonstrations from ARCTIC; pattern consistency (Pearson correlation between frame-to-frame human and predicted robot wrist displacements) exceeds 96% on all four structurally different hands, at 96.5% for Allegro and 99.1% for Shadow.

It designs a guidance predictor that combines the interaction pattern with a robot hand description (such as joint counts and link lengths and widths from a URDF) to estimate wrist motions and contact points adapted to the target hand. Unlike placing the robot wrist at the human wrist pose, it estimates wrist placement according to the target hand structure; unlike per-finger kinematic retargeting, it does not require finger-to-finger correspondence between human and robot hands. Wrist offsets track hand structure rather than overall size: Allegro and Paxini are placed about 5 cm behind the human wrist while Shadow and XHand are about 2.8 cm behind; contact precision (overlap between estimated and human contact points) is 74-80% across the four hands.

It trains a PPO policy guided by the estimated wrist motions and contact points, so the policy explores within what the target robot can execute and achieves higher success on bimanual manipulation of articulated objects. Compared with ObjDex, which guides the wrist with human wrist motions, adding target-hand-specific wrist placement and contact points stops the policy from imitating motions infeasible for that hand. Across 6 demonstrations with two Allegro hands on xArm6 arms, ObjDex averages 52.2% success, estimated wrist motions alone reach 58.1%, and wrist motions plus contact points reach 92.8%; each episode follows a 500-step reference trajectory and counts as success if position, rotation, and joint angle stay within 10 cm and the corresponding angles of the reference without hand-arm collision.

The interaction pattern is reusable across hand embodiments: freezing the pattern encoder and fine-tuning only the predictor and policy adapts to new hands, and policies trained in simulation transfer to a real robot. Prior methods typically retrain per robot hand or rely on teleoperation data; here simple fine-tuning suffices, and a new task needs only a single human demonstration video. After training with Allegro and freezing the pattern encoder, Shadow, Paxini, and XHand reach mean success rates of 84.0%, 73.5%, and 71.6%; in the real experiment CoTracker3 and HaMeR extract object and wrist trajectories from one video, and Allegro increases the microwave door angle over 100 steps, succeeding in 7 of 10 trials.

Perspective

The result targets learning bimanual dexterous manipulation of articulated objects from human-object demonstrations: inputs are object states (pose, articulation angle, shape) and human hand states plus a target robot hand description; outputs are wrist motions and contact points adapted to that hand, used to train a PPO policy. Evaluation runs in Isaac Gym, trained on 229 ARCTIC demonstrations and tested on 6 unseen ones covering Box, Espresso, Laptop, Microwave, Mixer, and Notebook; the real-robot validation uses an Allegro hand on an xArm6 opening a microwave. It is meant for researchers and engineering teams who want to obtain manipulation policies for different robot hands from few human demonstrations, especially when hand size or structure differs from the human hand and per-finger retargeting is infeasible.

Cross-hand transfer varies markedly on individual objects (Paxini reaches 12.6% on Mixer, XHand 37.3% on Notebook), suggesting transfer may depend on how well the object matches the hand, a dependence worth characterizing further. The real-robot validation covers only microwave opening over 10 trials, so behavior on more object categories and longer task horizons remains open. The predictor is also trained with wrist motions and contact points retargeted by an existing method as supervision, and how the quality of that supervision shapes results deserves more systematic study. Retargeting-based methods were not compared under the same reset protocol, so the 92.8% versus 52.2% gap should be read as a comparison against that specific baseline under the same initial-pose setting.

Sources