Having multiple applications in areas such as music curation, voice identification, and identifying environmental sounds, audio tagging proves to be a pivotal task within the realm of audio analysis and categorization. Due to their ability to capture complex characteristics from raw audio data, convolutional neural networks (CNNs) have found extensive utility in the context of audio tagging tasks. The selection of the feature extraction technique is an important component of CNN-based audio tagging. In this study, we investigate how three well-liked features such as MFCC, GFCC, and Spectrograms affect the precision of a CNN-based audio tagging model. Using a sizable dataset of labelled audio clips, we train a base CNN model with three convolutional layers and two fully connected layers, and we compare the model’s performance using each feature extraction technique. According to the outcomes of our experiments, Spectrograms perform better than MFCCs and GFCCs in terms of accuracy and F1 score. Finally, the performance of the model is evaluated on several performance matrices and this approach enables us to effectively classify audio files into various categories.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio Tagging – Deep Learning Approach

  • E. Sophiya,
  • S. Sudharsan,
  • S. K. Mukhil Varnan,
  • T. P. Nithishvaran,
  • M. S. Mohan Vamsi

摘要

Having multiple applications in areas such as music curation, voice identification, and identifying environmental sounds, audio tagging proves to be a pivotal task within the realm of audio analysis and categorization. Due to their ability to capture complex characteristics from raw audio data, convolutional neural networks (CNNs) have found extensive utility in the context of audio tagging tasks. The selection of the feature extraction technique is an important component of CNN-based audio tagging. In this study, we investigate how three well-liked features such as MFCC, GFCC, and Spectrograms affect the precision of a CNN-based audio tagging model. Using a sizable dataset of labelled audio clips, we train a base CNN model with three convolutional layers and two fully connected layers, and we compare the model’s performance using each feature extraction technique. According to the outcomes of our experiments, Spectrograms perform better than MFCCs and GFCCs in terms of accuracy and F1 score. Finally, the performance of the model is evaluated on several performance matrices and this approach enables us to effectively classify audio files into various categories.