ABSTRACT Micro‐expression analysis in long video sequences is a difficult problem owing to low‐resolution inputs, noise in the environment, inter‐subject variations, and the necessity of efficient temporal modeling of the subtle facial movements. To overcome these issues, this paper introduces a spatio‐temporal expression adaptive model (STEAM) that is a single framework of micro‐expression recognition that supports the localization of time. The suggested algorithm uses a window‐based temporal segmentation algorithm to extract the potential expression intervals in continuous video streams and then classify the intervals into emotion categories. The framework incorporates a Meta‐Attention Super‐Resolution (MASR) module to refine expression‐relevant parts of the face, a Robust Adaptive Noise Suppression (RANS) layer to reduce environmental distortions, and a lightweight Temporal Shift Module v2 (TSM‐v2) paired with transformer‐based encoding to capture subtle temporal motion patterns. Moreover, a Cross‐Domain Graph Attention Network (CD‐GAT) is employed to capture fine‐grained landmark‐level spatial relationships, and Adaptive Instance‐Specific Normalization (AISN) enhances the ability to deal with inter‐subject variability. In contrast to traditional micro‐expression temporal localization methods which assume the use of rigid event‐based temporal boundary detection based on temporal Intersection‐over‐Union (t‐IoU), the suggested framework implements a window‐based temporal localization algorithm, allowing to identify expression‐relevant intervals robustly and without the need to estimate the temporal boundaries. Demonstrations of high performance when using the Leave‐One‐Subject‐Out (LOSO) protocol on CASME I, CASME II and CAS(ME) 2 data in window‐level temporal localization and accurate emotion recognition are clear. These findings demonstrate the effectiveness of the suggested framework in the context of effective temporal analysis and subject‐independent micro‐expression recognition in the case of long video sequences.
Chauhan et al. (2026) studied this question.