Multimodal emotion recognition using residual self-attention and cross-modal fusion for audio-visual data | Synapse