Micro-expressions, subtle and often asymmetric facial movements, play a pivotal role in nonverbal emotional communication. Addressing the core challenges of temporal misalignment, fragmented feature extraction, and slow real-time detection in micro-expression recognition (MER), we propose a novel dual-branch spatiotemporal model for dynamic sequence MER. Leveraging MediaPipe for 3D facial feature extraction and Dynamic Time Warping (DTW) for sequence alignment, our method nonlinearly maps variable-length sequences to a fixed length. A hybrid data augmentation technique enhances model robustness, while the dual-branch network simultaneously captures local spatial features and global temporal dynamics. Experimental results on the CASMEII dataset demonstrate state-of-the-art performance with 99.22% accuracy, along with a significant improvement in real-time detection speed. This approach holds substantial practical value for applications in deception detection, mental health assessment, and human–computer interaction.
Yao et al. (Fri,) studied this question.