A Lightweight CNN-BiLSTM-Attention Framework for Speaker-Independent Speech Emotion Recognition Using CREMA_D Dataset

Authors

  • Amjed Kadhum Aljadiri
  • Mohammed S. H. Al-Tamimi

Keywords:

(Speech Emotion Recognition SER, Deep Learning DL, speaker independent, Feature Extraction, CREMA_D dataset)

Abstract

With applications in affective computing and human-computer interaction HCI, Speech Emotion Recognition (SER) seeks to deduce a speaker's emotional state from the acoustic information in speech.  However, due to the variability of speakers and the possibility of speaker overlap in the training and testing sets, it is difficult to measure the performance of deep learning-based SER systems. In this work, an efficient hybrid CNN-BiLSTM-Attention approach for the speech emotion recognition problem is proposed which uses the CREMA-D dataset. The proposed framework uses a low dimensional set of 42 handcrafted acoustic features which include MFCCs, delta MFCCs, delta MFCCs, pitch, zero crossing rate and RMS energy. The CNN component is learned to address local acoustic structure, and the BiLSTM models long-term temporal structure, along with the attention mechanism emphasizing emotionally informative speech frames. A strict speaker independent protocol is used to create a more realistic evaluation of model generalization, where the test set speakers are not used in training. The model is tested within six emotions: Angry, Disgust, Fear, Happy, Neutral, and Sad. The proposed framework obtains an accuracy of 71.72% at a test and a UAR of 71.95%. The result of per-emotion analysis shows that angry has the highest F1-score, which is 83.06%, while Neutral had the second highest F1-score, which is 79.09%, and Disgust is the most difficult class with F1 score of 59.20%. The results show that the proposed lightweight architecture could effectively realize the balanced emotion recognition and a more stringent assessment for generalization to speakers’ unseen during training.

Downloads

Published

2026-09-22

How to Cite

Aljadiri, A. K., & Al-Tamimi, M. S. H. (2026). A Lightweight CNN-BiLSTM-Attention Framework for Speaker-Independent Speech Emotion Recognition Using CREMA_D Dataset. International Journal of Artificial Intelligence and Machine Learning, 6(11s), 1045–1064. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/2247