AuraTracer智迹闻
中文

EVENT DOSSIER

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

The researchers proposed a new architecture called DSSM-CRF, aimed at addressing the challenge of coordinating acoustic evidence across time scales and bidirectional interaction processes in speech emotion recognition in dialogue scenarios. This architecture separates the impact of cross-speaker context from the internal emotional evolution of speakers: it uses a bidirectional state-space model to fuse self-supervised representations at the frame and dialogue levels, and the decoder sorts the utterances of each speaker into independent dynamic conditional random field chains. Auxiliary objective supervision tracks the emotional changes in each pair of consecutive utterances but does not participate in Viterbi inference, ensuring that the order of conversations affects emotional ratings without being considered as a transition in the trajectory of another speaker. On the IEMOCAP dataset, this method achieved an unweighted average accuracy of 75.81% and a weighted average accuracy of 74.90%; on the MELD dataset, it obtained a weighted average accuracy of 54.72% and an F1 score of 49.31%. Comparative experiments show that speaker factorization and CRF modeling lead to…

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
DSSM-CRFIEMOCAPMELD

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
DSSM-CRF × IEMOCAP1DSSM-CRF × MELD1IEMOCAP × MELD1

SignalsSIGNALS

Keyword heat
  • DSSM-CRF1
  • IEMOCAP1
  • MELD1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

研究人员提出 DSSM-CRF 架构,用于解决对话中语音情感识别需协调跨时间尺度声学证据与双向交互过程的问题。该音频仅架构将跨说话人上下文影响与说话人内情感演变分离:双向状态空间模型在帧级和对话级编码融合自监督表示,解码器将各说话人 utterance 排序为独立动态条件随机场链。辅助目标监督每对连续 utterance 的情感变化但不参与 Viterbi 推理,确保互话轮次影响上下文情感评分而不被视为另一说话人的轨迹过渡。在 IEMOCAP 数据集上,该方法实现 75.81% UA 和 74.90% WA;在 MELD 数据集上取得 54.72% WA 和 49.31% WF1。匹配对照实验显示说话人因子分解与 CRF 建模带来互补增益。