Automatic recognition of emotions from audio data is a growing field with applications in various domains. However, challenges exist regarding pre-processing complexities, language dependence, and the need for more robust models. This paper addresses these issues by proposing a framework for audio emotion recognition. Our methodology involves thorough data pre-processing, including analysis of class distribution, feature extraction using spectrograms, MFCCs, and time-domain characteristics, and data balancing techniques. The paper investigates a Multi-Model approach incorporating a blend of deep learning and machine learning algorithms, which include random forest regressor, gradient boosting, Support vector machine, K-neighbors classifier, MLP classifier, XGB classifier, and deep learning techniques using a custom CNN model. We selected the custom CNN model due to its ability to capture intricate temporal and spectral features inherent in audio data, enhancing emotion recognition with an accuracy of 96.34%. The results demonstrate the effectiveness of our proposed framework for audio emotion recognition. We identify dataset diversity and linguistic coverage limitations, highlighting areas for future research. Overall, this work contributes valuable insights into the current state of audio emotion recognition and paves the way for further advancements in this field.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unveiling Emotions from Audio: A Multi-model Exploration Leveraging Diverse Datasets

  • Shashank Mouli Satapathy,
  • Vaibhav Vilas Pawar,
  • Atharva Gulkotwar

摘要

Automatic recognition of emotions from audio data is a growing field with applications in various domains. However, challenges exist regarding pre-processing complexities, language dependence, and the need for more robust models. This paper addresses these issues by proposing a framework for audio emotion recognition. Our methodology involves thorough data pre-processing, including analysis of class distribution, feature extraction using spectrograms, MFCCs, and time-domain characteristics, and data balancing techniques. The paper investigates a Multi-Model approach incorporating a blend of deep learning and machine learning algorithms, which include random forest regressor, gradient boosting, Support vector machine, K-neighbors classifier, MLP classifier, XGB classifier, and deep learning techniques using a custom CNN model. We selected the custom CNN model due to its ability to capture intricate temporal and spectral features inherent in audio data, enhancing emotion recognition with an accuracy of 96.34%. The results demonstrate the effectiveness of our proposed framework for audio emotion recognition. We identify dataset diversity and linguistic coverage limitations, highlighting areas for future research. Overall, this work contributes valuable insights into the current state of audio emotion recognition and paves the way for further advancements in this field.