Unveiling Emotions from Audio: A Multi-model Exploration Leveraging Diverse Datasets
摘要
Automatic recognition of emotions from audio data is a growing field with applications in various domains. However, challenges exist regarding pre-processing complexities, language dependence, and the need for more robust models. This paper addresses these issues by proposing a framework for audio emotion recognition. Our methodology involves thorough data pre-processing, including analysis of class distribution, feature extraction using spectrograms, MFCCs, and time-domain characteristics, and data balancing techniques. The paper investigates a Multi-Model approach incorporating a blend of deep learning and machine learning algorithms, which include random forest regressor, gradient boosting, Support vector machine, K-neighbors classifier, MLP classifier, XGB classifier, and deep learning techniques using a custom CNN model. We selected the custom CNN model due to its ability to capture intricate temporal and spectral features inherent in audio data, enhancing emotion recognition with an accuracy of 96.34%. The results demonstrate the effectiveness of our proposed framework for audio emotion recognition. We identify dataset diversity and linguistic coverage limitations, highlighting areas for future research. Overall, this work contributes valuable insights into the current state of audio emotion recognition and paves the way for further advancements in this field.