Comparative Analysis of Speech Emotion Recognition Models
摘要
The field of Speech Emotion Recognition (SER) has accumulated significant traction in recent years, driven by its potential applications in human-computer interaction, medical, and education. This study investigates the effectiveness of various SER models by conducting a comparative analysis using a specially curated dataset. The dataset comprises voice recordings from 129 female participants, each expressing seven distinct emotions while uttering the neutral sentence “The cat is sleeping”. This design gives SER models a richer and more realistic evaluation platform by allowing the collection of subtle emotional fluctuations within an impartial setting. Five well-known machine learning models are used and examined in this study. Each model’s performance is thoroughly assessed using important metrics including accuracy and F1 score, and the results are displayed using loss graphs and bar plots. A greater comprehension of the merits and demerits of each model in the context of SER is made possible by this thorough examination. The aim of this research is to enhance SER technology and its applications by comparing SER models, identifying research gaps, and suggesting future paths.