Multi-Modal Demonstrations for Dexterous Manipulation
Multi-Modal Demonstrations for Dexterous Manipulation
- 1 Keio University
- 2 University of California, Berkeley
- 3 Carnegie Mellon University
Abstract
Collecting demonstrations for long-horizon, contact-rich manipulation remains a major bottleneck in robot learning. Teleoperation scales to long-horizon tasks but is imprecise for fine, contact-rich interaction, whereas kinesthetic teaching excels at contact-rich manipulation but is physically exhausting and, critically, records the demonstrator’s hands in every frame. This paper investigates how teleoperation and kinesthetic teaching can be combined for training policies for dexterous manipulation. We propose a demonstration framework that uses teleoperation data for broad task execution and kinesthetic teaching for difficult fine-control phases. We first train a diffusion policy on teleoperated demonstrations, then refine it through real-time kinesthetic alignment. Because the human hand appears during correction but not at deployment, we supervise the visual encoder with object masks so it attends to the manipulated object rather than the demonstrator’s hand. We were able to visualize the visual encoder that attends more to the manipulated object. This work suggests that combining heterogeneous demonstration types allows a single policy to handle both long-horizon execution and fine-grained contact-rich phases of dexterous manipulation.