PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 17, 2026Concurrency and Computation Practice and Experience0 citations

Enhanced Spatial–Temporal Transformer Network for Robust Micro‐Expression Temporal Localization and Recognition in Long Video Sequences

View Full Paper
ACAnjaly ChauhanACAbhishek ChaudharySJShikha Jain

Key Points

  • The aim is to develop a robust framework for micro-expression recognition and temporal localization in long video sequences.
  • Introduced a spatio-temporal expression adaptive model (STEAM) for classification of emotion categories.
  • Implemented window-based temporal segmentation and a Meta-Attention Super-Resolution (MASR) module to refine facial expression.
  • Utilized a Cross-Domain Graph Attention Network (CD-GAT) to capture fine-grained spatial relationships.
  • Achieved high performance in window-level temporal localization with strong accuracy in emotion recognition.
  • Performed well when evaluated against CASME I, CASME II, and CAS(ME) 2 datasets using Leave-One-Subject-Out (LOSO) protocol.

Abstract

ABSTRACT Micro‐expression analysis in long video sequences is a difficult problem owing to low‐resolution inputs, noise in the environment, inter‐subject variations, and the necessity of efficient temporal modeling of the subtle facial movements. To overcome these issues, this paper introduces a spatio‐temporal expression adaptive model (STEAM) that is a single framework of micro‐expression recognition that supports the localization of time. The suggested algorithm uses a window‐based temporal segmentation algorithm to extract the potential expression intervals in continuous video streams and then classify the intervals into emotion categories. The framework incorporates a Meta‐Attention Super‐Resolution (MASR) module to refine expression‐relevant parts of the face, a Robust Adaptive Noise Suppression (RANS) layer to reduce environmental distortions, and a lightweight Temporal Shift Module v2 (TSM‐v2) paired with transformer‐based encoding to capture subtle temporal motion patterns. Moreover, a Cross‐Domain Graph Attention Network (CD‐GAT) is employed to capture fine‐grained landmark‐level spatial relationships, and Adaptive Instance‐Specific Normalization (AISN) enhances the ability to deal with inter‐subject variability. In contrast to traditional micro‐expression temporal localization methods which assume the use of rigid event‐based temporal boundary detection based on temporal Intersection‐over‐Union (t‐IoU), the suggested framework implements a window‐based temporal localization algorithm, allowing to identify expression‐relevant intervals robustly and without the need to estimate the temporal boundaries. Demonstrations of high performance when using the Leave‐One‐Subject‐Out (LOSO) protocol on CASME I, CASME II and CAS(ME) 2 data in window‐level temporal localization and accurate emotion recognition are clear. These findings demonstrate the effectiveness of the suggested framework in the context of effective temporal analysis and subject‐independent micro‐expression recognition in the case of long video sequences.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chauhan et al. (2026) studied this question.

synapsesocial.com/papers/6a095ba67880e6d24efe17afhttps://doi.org/10.1002/cpe.70752
Ask AI
Helpful
Bookmark
Share
View Full Paper