AuraTracer智迹闻
中文

EVENT DOSSIER

Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated

The researchers proposed a new multi-modal emotion recognition framework that combines audio and video data for emotion analysis. The framework extracts semantic embeddings, MFCC features, and acoustic statistics at the audio level, and uses BiLSTM for alignment and fusion; at the video level, the ResNet50-BiLSTM architecture is used to extract spatiotemporal features. Additionally, a feature-level fusion mechanism based on多头 attention is introduced to enhance the synergistic effect between modalities. Experiments show that this model achieves significantly better accuracy and robustness on the MELD and IEMOCAP datasets, especially in scenarios with data imbalance.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
BiLSTMIEMOCAPMELDResNet50Wav2Vec2

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
BiLSTM × IEMOCAP1BiLSTM × MELD1BiLSTM × ResNet501BiLSTM × Wav2Vec21IEMOCAP × MELD1IEMOCAP × ResNet501

SignalsSIGNALS

Keyword heat
  • Wav2Vec21
  • ResNet501
  • BiLSTM1
  • MELD1
  • IEMOCAP1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion

本文提出了一种基于多特征编码与注意力融合的多模态情感识别新框架。该框架在音频端提取语义嵌入、MFCC 特征及声学统计量,通过 BiLSTM 对齐融合;在视频端采用 ResNet50-BiLSTM 架构提取时空特征;并引入基于多头注意力的特征级融合机制以增强模态协同。在 MELD 和 IEMOCAP 数据集上的实验表明,该模型在准确性和鲁棒性上显著优于基线方法,且注意力融合策略在数据不平衡场景下表现更优。