DexTaG is an efficient RL pipeline for learning dexterous tool-use policies from mocap-glove demonstrations.
Its key idea is to use the glove's tactile readouts as reward guidance during RL training, which guides contact acquisition and human-like grasp formation while reducing reliance on precise reference geometry. The trained retargeter is then distilled into a tactile-free student controller that deploys zero-shot in the real world.
Mocap gloves are a promising path for scaling dexterous manipulation data: they capture dexterous finger motion while the demonstrator manipulates the object directly. The recorded motions are then retargeted to bridge the kinematic gap between the human and robot hand. We use a commercial EMF glove with a full-hand tactile sensor array and process each capture into a hand–object trajectory used as reference for retargeter training.
Bare-handed manipulation with the glove unlocks tasks that require in-hand reorientation.
Processed hand–object trajectory
We collect two contact-rich tool-use tasks—marker pen and hammer—both of which require non-trivial in-hand reorientation.
For each task, we train a single retargeting policy over many reference trajectories of the same object. Sharing experience across trajectories helps the policy generalize to held-out references within the same task, without per-trajectory retraining. The glove's tactile readouts provide reward guidance for contact acquisition and human-like grasp formation, reducing reliance on precise reference geometry.
Rollouts of our generalizable retargeter on held-out trajectories. Each panel shows the reference (left), the policy rollout (middle), and the target vs. real-time simulated tactile maps used to compute the tactile rewards (right).
The retargeter is distilled into a tactile-free student controller whose reference input contains only the target object trajectory, with no reference hand poses. The controller uses visual–proprioceptive feedback for real-world deployment. In simulation, the student retains generalization to held-out object trajectories within the same object and task setting. Below are real-world rollouts on collected references: the left panel is the reference object trajectory, and the right panel is the rollout played at 1× real-time speed.
Because the student's reference input contains only the target object trajectory, we can synthesize new trajectories for it to follow using simple rule-based procedures. Below we show qualitative examples of synthesized pen-writing-letter trajectories in simulation (top) and the real world (bottom): the left panel is the synthesized trajectory, and the right panel is the rollout. Real-world rollouts are played at 1× real-time speed.
Each video shows three panels, from left to right: Reference, Ours, and Geom-gated. Geom-gated uses the same fingertip-proximity reward but replaces measured tactile contact gates with distance-based labels from the motion reference and omits tactile-map similarity. In these examples, tactile-as-guidance produces a more natural marker-pen grasp and enables hammer pickup, where Geom-gated struggles.
Representative failure cases observed during real-world deployment.
Failed to pick up the marker pen.
Failed to reorient the marker in-hand into a valid manipulation pose.
Failed to place the marker onto the table.
Sub-optimal grasp pose due to a sim-to-real gap: the real-world friction between the fingers is higher than in simulation, leaving the thumb floating rather than supporting the marker.
Failed to pick up and lift the hammer.
Failed to place the hammer onto the table, also due mainly to the sim-to-real gap: contact of the hand's soft rubber cover is hard to simulate faithfully.
Self-collision with the wrist camera, whose collision volume was not modeled during simulation training.