Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
The researchers proposed a new architecture called DSSM-CRF, aimed at addressing the challenge of coordinating acoustic evidence across time scales and bidirectional interaction processes in speech emotion recognition in dialogue scenarios. This architecture separates the impact of cross-speaker context from the internal emotional evolution of speakers: it uses a bidirectional state-space model to fuse self-supervised representations at the frame and dialogue levels, and the decoder sorts the utterances of each speaker into independent dynamic conditional random field chains. Auxiliary objective supervision tracks the emotional changes in each pair of consecutive utterances but does not participate in Viterbi inference, ensuring that the order of conversations affects emotional ratings without being considered as a transition in the trajectory of another speaker. On the IEMOCAP dataset, this method achieved an unweighted average accuracy of 75.81% and a weighted average accuracy of 74.90%; on the MELD dataset, it obtained a weighted average accuracy of 54.72% and an F1 score of 49.31%. Comparative experiments show that speaker factorization and CRF modeling lead to…