ABSTRACT To address the limitations of insufficient electroencephalogram (EEG) feature extraction and instability in cross‐modal decision fusion, which reduce the robustness of existing multimodal sentiment recognition systems, we propose a structured fusion‐based multimodal sentiment recognition (SF‐MER) method. SF‐MER separately extracts EEG and facial expression features and classifies them using independent classifiers. A modality‐quality‐aware Dempster–Shafer (D–S) decision fusion strategy is subsequently employed to fuse the classification results and generate the final emotion category. For EEG signals, we propose a graph‐based attention‐enhanced network to model spatiotemporal and spectral dependencies. The EEG branch employs FCM–GC–PLI fusion features to extract rich multichannel EEG details, constructing Brain‐Region Frequency‐Band Cross‐Attention (BFCA) to enhance the model's focus on sentiment‐correlated features. The facial expression branch uses depth‐separable residual shrink convolution (DSRS‐CN) to extract compact and robust visual features and employs average‐maximum dual pooling to filter sentimental peak frames, reducing redundancy. Finally, a modality‐quality‐aware decision fusion strategy outputs a sentiment recognition result. This approach achieved average recognition rates of 98.50% and 98.36% on the arousal and valence dimensions of the DEAP and MAHNOB‐HCI datasets, respectively, representing improvements of 1.77% and 1.58% over the baseline model. Experimental results demonstrate that the proposed method outperforms existing approaches.
Liu et al. (Thu,) studied this question.