Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated
The researchers proposed a new multi-modal emotion recognition framework that combines audio and video data for emotion analysis. The framework extracts semantic embeddings, MFCC features, and acoustic statistics at the audio level, and uses BiLSTM for alignment and fusion; at the video level, the ResNet50-BiLSTM architecture is used to extract spatiotemporal features. Additionally, a feature-level fusion mechanism based on多头 attention is introduced to enhance the synergistic effect between modalities. Experiments show that this model achieves significantly better accuracy and robustness on the MELD and IEMOCAP datasets, especially in scenarios with data imbalance.