PatternDex:
Learning Interaction Patterns to Guide Reinforcement Learning of Bimanual Dexterous Manipulation of Articulated Objects

David Minkwan Kim, Runfa Blark Li, Beckham Po-Ju Lee, Nikolay Atanasov, Truong Nguyen
University of California, San Diego
Overview of PatternDex

Abstract

In this paper, we develop a method that enables bimanual dexterous hands to manipulate articulated objects with a high success rate without suffering from an embodiment gap. We observe that the correlation between hand motions and object motions is dictated by the object rather than the hands and can be learned from human-object demonstrations. Based on this observation, we propose PatternDex, a method that learns this correlation and represents it as a token sequence, which we call an interaction pattern. From this pattern, PatternDex estimates the wrist motions and contact points that fit the target robot, and then trains a reinforcement learning policy that exploits these estimates as guidance. Since the guidance fits the target embodiment, the policy explores only the actions that the target robot can execute and thus achieves high success rates. PatternDex also requires only simple fine-tuning to train a new robot, since it can reuse the learned interaction pattern. We evaluate PatternDex with bimanual dexterous hands on human demonstrations from the ARCTIC dataset. PatternDex achieves, on average, a 92.8% success rate with Allegro hands, while the state-of-the-art baseline achieves 52.2%. Also, it achieves success rates above 70% with three other robot hands after fine-tuning alone. Furthermore, we verify that the learned policy transfers well to a real-world task of opening a microwave.

Video

Method


Overview of PatternDex model

PatternDex model

1. Interaction Pattern Encoder

A spatial network \(f\) (MLP) encodes the object state and the human wrist pose at each frame into a spatial token that describes the relative pose of the wrists with respect to the object. A temporal network \(\mathcal{T}\) (4 Transformer layers, 4 heads) stacks these tokens over a chunk of \(H+1 = 64\) frames and learns how they influence one another. The final hidden state is the interaction pattern \(P_I\), which contains information independent of the human hand and can therefore be shared across embodiments.

2. Interaction Pattern Decoder

An embodiment encoder \(\phi\) turns the robot hand description (number of joints, link lengths and widths of each finger and the palm, parsed from URDF) into embodiment tokens, which are concatenated with \(P_I\) as memory \(\mathcal{M}\). Following ACT-style chunk-wise prediction, a 6-layer Transformer decoder \(\mathcal{D}\) with learnable queries \(Q\) cross-attends to this memory and outputs, through a wrist head and a contact head, the robot wrist motions and the contact points for the whole chunk at once. For a new robot hand, only \(\mathcal{D}\) and \(\phi\) are fine-tuned while the encoder stays frozen.

3. Guided Policy Learning

A PPO policy \(\pi_\theta\) receives the current state, the reference object trajectory, and the estimated contact points, and outputs the finger joint targets of both hands together with a residual correction on top of the estimated wrist motions. The reward is the sum of an object reward, which measures how closely the manipulated object tracks the reference in position, rotation, and articulation angle, and a contact reward, which rewards fingertips and palm for covering the estimated contact points and penalizes fingertips that touch the object far from them.

Interaction Pattern

Human demonstration (ARCTIC)

Interaction pattern: encode and decode

Interaction pattern for opening the microwave

Estimated Allegro wrist motions and contact points

Interaction pattern for opening the microwave. The object trajectory is shown with its articulation angle only (dashed orange). The human wrist motions (red for the left hand, blue for the right hand) change as the articulation angle increases, and the interaction pattern encodes this change. The light red and light blue curves are Allegro wrist motions estimated from the interaction pattern. They preserve the shape of the human wrist motions and are shifted back according to the size of Allegro, which shows that the estimated wrist motions are adjusted to the robot.

Simulation Results


Allegro Hand Allegro Hand

Box SR 99.9%

Espresso machine SR 98.7%

Laptop SR 94.2%

Microwave SR 98.6%

Mixer SR 99.5%

Notebook SR 66.0%


Cross-Embodiments

Shadow Hand Shadow Hand

Paxini Hand Paxini Hand

XHand XHand


Quantitative Results

Method Object / Traj ID SR (%) ↑ Pos. err. (cm) ↓ Rot. err. (°) ↓ Joint err. (°) ↓ Contact err. (cm) ↓
ObjDexBox / s04_0150.11.011.586.14–
PatternDex (w/o Contact)87.41.732.211.99–
PatternDex (full)99.92.053.021.922.35
ObjDexEspresso / s01_0145.12.164.132.08–
PatternDex (w/o Contact)47.82.876.749.87–
PatternDex (full)98.71.994.401.943.61
ObjDexLaptop / s09_0277.40.521.145.92–
PatternDex (w/o Contact)51.32.931.723.23–
PatternDex (full)94.22.243.132.194.24
ObjDexMicrowave / s04_0198.43.131.712.58–
PatternDex (w/o Contact)76.23.282.562.98–
PatternDex (full)98.62.422.401.821.92
ObjDexMixer / s08_0134.41.763.503.36–
PatternDex (w/o Contact)73.22.254.195.16–
PatternDex (full)99.52.773.912.412.91
ObjDexNotebook / s08_037.80.702.085.79–
PatternDex (w/o Contact)12.52.274.793.25–
PatternDex (full)66.02.944.684.066.65
ObjDexMean52.21.552.364.31–
PatternDex (w/o Contact)58.12.553.704.41–
PatternDex (full)92.82.403.592.393.61
Quantitative results of the guided policy on the ARCTIC dataset. SR denotes success rate. Pos., Rot., and Joint err. denote the position, rotation, and joint angle errors of the object. Contact err. denotes the average distance from the robot hand to the estimated contact points. Standard deviations are reported in the paper.

Robot Box / s04_01 Espresso / s01_01 Laptop / s09_02 Microwave / s04_01 Mixer / s08_01 Notebook / s08_03 Mean
Allegro99.9±0.198.7±1.394.2±5.198.6±2.899.5±0.766.0±2.292.8±12.2
Shadow96.3±4.588.5±5.651.6±2.378.0±3.297.0±2.092.3±4.784.0±15.8
Paxini96.7±3.274.9±2.993.8±5.899.0±1.312.6±0.064.0±1.773.5±30.0
XHand95.4±5.073.8±3.740.0±6.791.9±7.091.0±3.937.3±1.871.6±24.3
Cross-embodiment success rates (%) of the guided policy. The pattern encoder \(\mathcal{E}\) is trained with Allegro and kept frozen, and only the decoder \(\mathcal{D}\) and the embodiment encoder \(\phi\) are fine-tuned for Shadow, Paxini, and XHand.

Sim-to-Real: One Human Video Is Enough

Since PatternDex has already learned the interaction pattern from ARCTIC, a new real-world task needs only a single human demonstration video and no robot demonstrations. CoTracker3 tracks the position and the hinge angle of the microwave, HaMeR tracks the human wrist pose, and the tracked trajectories are fed into the guidance predictor. The guided policy is trained in Isaac Gym and deployed directly on an Allegro hand mounted on an xArm6. The robot opens the door in 7 of 10 trials, with the door angle increasing from 0° to 74.2° over 100 steps.


Pipeline

Tracked wrist and object state → guidance → simulator training → real robot

Sim-to-real pipeline

Real-world rollout

Allegro hand on xArm6 opening the microwave with the policy trained in simulation

The insets show the tracked door angle, the front camera view, and the human demonstration used as reference.

BibTeX

@article{kim2026patterndex,
  title   = {PatternDex: Learning Interaction Patterns to Guide Reinforcement Learning of Bimanual Dexterous Manipulation of Articulated Objects},
  author  = {Kim, David Minkwan and Li, Runfa Blark and Lee, Beckham Po-Ju and Atanasov, Nikolay and Nguyen, Truong},
  journal = {arXiv preprint arXiv:2610.04765},
  year    = {2026}
}