Back close

Hybrid GraphSENet: An efficient hybrid graph network for Speech Emotion Recognition

Publication Type : Journal Article

Publisher : Elsevier BV

Source : Computers and Electrical Engineering

Url : https://doi.org/10.1016/j.compeleceng.2026.111462

Keywords : SER, MFCC, Visibility graphs, CNN, Graph features, Hybrid model, DL, Heatmaps

Campus : Amritapuri

School : School of Engineering

Department : Electronics and Communication

Year : 2026

Abstract : Speech Emotion Recognition (SER) has gained increasing interest over the past few years, with most approaches relying on features obtained from time, frequency, or cepstral domains. Although these approaches are proven to be effective, they often treat these features independently, failing to capture the underlying deep structural and temporal relationships. The proposed work uses a novel Mel Frequency Cepstral Coefficient-Natural Visibility Graph (MFCC-NVG) algorithm, which converts audio-based Mel Frequency Cepstral Coefficients into complex networks employing Visibility Graphs. We propose a Hybrid GraphSENet architecture that integrates hand-crafted graph-based features with deep learned features obtained from heatmap representations of the adjacency matrix, derived using the MFCC-NVG algorithm. We concatenate these features and feed them to a Deep Neural Network (DNN) to assess their joint performance on SER. We also experiment on standalone SER models, viz., graph-based hand-crafted features with baseline learning models, and heatmap representations of adjacency matrices with Deep Learning models. We validate our proposed model in both speaker-dependent and speaker-independent scenarios, employing three datasets: Emo-DB, IEMOCAP, and Amritaemo_Arabic. The Hybrid GraphSENet architecture consistently exhibits superior performance, generating an average SER accuracy of 96.83%, 77.53% and 97.89% for the speaker-dependent scenario and, 73.90%, 67.65% and 72.37% for the speaker-independent scenario, employing Emo-DB, IEMOCAP and Amritaemo_Arabic datasets, respectively, effectively outperforming several state-of-the-art models.

Cite this Research Publication : Sreedutt Ram J., Poorna S.S., Anuraj K., Hybrid GraphSENet: An efficient hybrid graph network for Speech Emotion Recognition, Computers and Electrical Engineering, Elsevier BV, 2026, https://doi.org/10.1016/j.compeleceng.2026.111462

Admissions Apply Now