PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 16, 2026Science Robotics6 citations

Learning realistic lip motions for humanoid face robots

View Full Paper
YHY. Charlie HuJLJiong LinJGJudah Allen Goldfeder

Key Points

  • The aim is to develop a humanoid robot face that achieves realistic lip motions synchronized with speech.
  • Developed a humanoid robot with soft silicone lips actuated by a 10-degree-of-freedom mechanism.
  • Implemented a self-supervised learning pipeline using variational autoencoder and facial action transformer.
  • Enabled the robot to infer lip trajectories directly from speech audio without predefined movements.
  • The new method shows superior performance in lip-audio synchronization compared to traditional amplitude-based heuristics.
  • Achieved visually coherent lip motions, enhancing the realism of robot speech in various linguistic contexts.
  • Generalized successfully across 10 languages not used during the training phase.

Abstract

Lip motion represents outsized importance in human communication, capturing nearly half of our visual attention during conversation. Yet anthropomorphic robots often fail to achieve lip-audio synchronization, resulting in clumsy and lifeless lip behaviors. Two fundamental barriers underlay this challenge. First, robotic lips typically lack the mechanical complexity required to reproduce nuanced human mouth movements; second, existing synchronization methods depend on manually predefined movements and rules, restricting adaptability and realism. Here, we present a humanoid robot face designed to overcome these limitations, featuring soft silicone lips actuated by a 10–degree-of-freedom mechanism. To achieve lip synchronization without predefined movements, we used a self-supervised learning pipeline based on a variational autoencoder (VAE) combined with a facial action transformer, enabling the robot to autonomously infer more realistic lip trajectories directly from speech audio. Our experimental results suggest that this method outperforms simple heuristics like amplitude-based baselines in achieving more visually coherent lip-audio synchronization. Furthermore, the learned synchronization successfully generalizes across multiple linguistic contexts, enabling robot speech articulation in 10 languages unseen during training.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2026) studied this question.

synapsesocial.com/papers/6969d428940543b9777091c7https://doi.org/10.1126/scirobotics.adx3017
Ask AI
Helpful
Bookmark
Share
View Full Paper