PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 15, 2026Actuators0 citationsOpen Access

HAML: Humanoid Adversarial Multi-Skill Learning via a Single Policy

View Full Paper
XFXu FangHLHsin‐Yi LiaoYCYanyun Chen

Key Points

  • The central aim is to develop a humanoid control system that efficiently translates motion datasets into a single-policy multi-skill framework.
  • Developed a two-stage learning system mapping motion datasets to a humanoid controller.
  • Utilized one-hot skill labels for automatic dataset construction with minimal manual effort.
  • Implemented a condition-aware loss to enhance controllability and reduce mode collapse.
  • Employed a teacher-student policy distillation strategy for robust real-world application.
  • Improved skill and transition coverage compared to previous approaches.
  • Enhanced realism and training efficiency in humanoid motion control.
  • Successfully operated on real hardware at 100 Hz with low latency of 15–25 ms.

Abstract

Translating large-scale motion datasets into robust, deployable humanoid controllers is a critical challenge in engineering informatics, primarily due to the scarcity of high-quality annotations, the risk of mode collapse in conditional generation, and the strict constraints of onboard computing hardware. This paper presents a deployable two-stage learning system that maps clip-level motion datasets to a single-policy multi-skill controller and its deployable counterpart. We adopt coarse one-hot skill labels that can be assigned automatically at the clip level with negligible manual effort, enabling scalable dataset construction. To prevent conditional discriminators from ignoring skill conditions, we inject mismatched (transition, label) pairs and introduce a condition-aware loss that explicitly penalizes incorrect transition–label associations, improving controllability and mitigating mode collapse. For real-world deployment, we further propose a two-stage training strategy: a privileged teacher policy is first trained in simulation and then distilled into a student policy that relies on stacked historical proprioceptive observations, ensuring robustness against sensing noise and latency without relying on external state estimation. Extensive evaluations in simulation and on real hardware demonstrate improved skill coverage, transition coverage, realism, and training efficiency across heterogeneous embodiments. With the onboard computer of a Unitree G1 robot, the distilled policy runs at 100 Hz with 15–25 ms latency, confirming the system’s engineering feasibility.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fang et al. (2026) studied this question.

synapsesocial.com/papers/69df2c2fe4eeef8a2a6b131bhttps://doi.org/10.3390/act15040212
Ask AI
Helpful
Bookmark
Share
View Full Paper