Accurate perception and replication of dynamic human motion represent a core challenge in advancing the autonomy and naturalistic movement of humanoid robots. This research, aligned with the goals of humanoid robotics, investigates this challenge through the domain of Wushu Sanda—a complex martial art with semantically rich and varied actions. We propose a novel action recognition framework based on the semantic feature matching of salient motion images to bridge perception and imitation. Human kinematic features are first extracted via joint distance and angle calculations to form a spatiotemporal feature map. Singular Value Decomposition (SVD) is employed for data compression and redundancy reduction, facilitating the evaluation of pixel significance within salient regions for precise semantic matching. This process feeds into an optimized Deep Convolutional Neural Network (CNN) to construct a robust recognition model for Wushu Sanda movements. Experimental validation confirms the method' s efficacy, demonstrating efficient extraction of discriminative features with a maximum Gini index of 0.10, stable recognition loss converging near 0.1, and an F1-score consistently approaching 1.0 across different sample sizes. The results substantiate the proposed method' s effectiveness and underscore its potential for enhancing motion learning and imitation capabilities in humanoid robotic systems.
Shan Li (Thu,) studied this question.