University of Pennsylvania · Human-Robot Interaction · Speech-to-ASL

Voice-Controlled Robotic Fingerspelling System

A real-time HRI pipeline that translates spoken English into ASL fingerspelling on a simulated 24-DoF DexHand — speech recognition, logic parsing and trajectory execution wired together in ROS 2, with poses retargeted from video of a human hand.

Period

Oct – Nov 2025

Stack

ROS 2 Humble · Gazebo · MediaPipe · OpenCV · ros2_control · Flask

Highlights

  • 01Modular ROS 2 Humble architecture chaining speech recognition, logic parsing and trajectory execution in Gazebo
  • 02Custom Python retargeting tool built on OpenCV and MediaPipe, mapping 3D human hand landmarks to normalized robot joint angles — 10x faster pose calibration
  • 03Extended the robot's ros2_control interface for dynamic trajectory execution across 26+ custom-tuned alphabet poses
  • 04Worked around WSL2 hardware isolation with a low-latency Flask TCP/IP client-server bridge streaming audio and video from Windows into the Linux ROS environment
  • 05Demonstrated end to end by fingerspelling whole words — the demo spells FLASH

Write-up

The goal was a robot hand that spells what it hears, in real time, without a human hand-authoring every letter pose. The architecture is deliberately modular: a speech recognition stage, a parsing stage that turns recognized text into a letter stream, and an execution stage that drives a 24-DoF DexHand in Gazebo — three ROS 2 Humble nodes that can each be swapped or debugged in isolation.

The bottleneck was never the speech, it was calibration. Hand-tuning 24 joints for each of 26 letters is slow, and the poses it produces read as mechanical rather than as ASL. So instead of authoring them I built a retargeting tool in OpenCV and MediaPipe that takes 3D hand landmarks from video of a person signing and maps them onto the robot's kinematic chain as normalized joint angles. That made pose calibration roughly ten times faster, and the poses far more legible as actual fingerspelling.

The problem I did not expect to spend time on was the operating system boundary. ROS 2 runs in WSL2, and WSL2 does not give the Linux side direct access to the microphone or camera — so the two things the pipeline is built around were both on the wrong side of the boundary. Rather than move the stack to bare metal, I wrote a small Flask client-server bridge that streams audio and video over TCP/IP from Windows into the Linux environment with low enough latency to keep the interaction feeling real-time.

Execution runs through the robot's ros2_control interface, extended to accept dynamic trajectories so letters can be sequenced into words on the fly rather than replayed from a fixed script.

Links