VERG: HYBRID MACHINE LEARNING BASED SPEECH EMOTION RECOGNITION SYSTEM
DOI:
https://doi.org/10.4238/8kv1tf14Keywords:
health literacy, cardiovascular disease, concept paper, knowledge gap, adult learning, intervention design, teach-back, plain language, community health worker, health equityAbstract
This research delineates a hybrid framework for Voice-based Emotion Recognition and Gender Classification (VERG) employing an ensemble approach that integrates Deep Learning (DL) and Machine Learning (ML) techniques. Two benchmark datasets were employed ; The Toronto Emotional Speech Set (TESS) comprises 2800 samples for emotion recognition, whereas the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains an initial 1440 samples for emotion and gender analysis. To mitigate the data scarcity problem in RAVDESS and improve generalization, data augmentation via SpecAugment-based data augmentation was implemented to expand the dataset to 5760 samples and to strengthen the model’s resilience to voice variances. Features were derived utilizing Mel-Frequency Cepstral Coefficients (MFCCs). The framework was evaluated using three models: Random Forest (RF), Long Short-Term Memory (LSTM), and a stacking ensemble of both. The models exhibited exceptional performance on TESS, achieving 98.86% for LSTM, 99.43% for RF, and 99.71% for stacking. In the RAVDESS emotion-only classification, the Random Forest (RF) model attained an accuracy of 84.14%, while the Long Short-Term Memory (LSTM) model achieved 82.06%. The accuracy was enhanced to 87.27% with the application of stacking. The models exhibited strong performance in joint emotion and gender classification on the enlarged RAVDESS dataset (5760 samples), achieving accuracy rates of 86.10% with RF, 84.70% with LSTM, and 88.50% with stacking. The overall results affirm the competitiveness of RF and LSTM, while also demonstrating the stacking ensemble's superior accuracy, underscoring the efficacy of hybrid methodologies in capturing the temporal and statistical attributes of speech for practical applications.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

