In the rapidly evolving field of audio analysis, distinguishing between diverse audio formats like Podcasts (Talk), Advertisements and Songs poses unique challenges due to their varying content, structure, and features. Addressing this issue, our paper, introduces a novel approach to classifying these distinct types of audio content through advanced deep learning techniques. To precisely understand the core attributes of each audio format, we deployed a thorough extraction of features through two methods: the first being Mel-Frequency Cepstral Coefficients (MFCCs), and the second, the application of decibel scaling to the Mel spectrogram feature, which transform audio information into spectrograms. Utilizing these spectrograms as a foundation, our study delves into the training of various Convolutional Neural Network (CNN) frameworks, including VGGNet, DenseNet, and MobileNet. We have collected Audio data from various open-source platforms, including Google Podcasts, YouTube, among others. Initially, we developed binary classification models to categorize Hindi and English audio content into three distinct pairs: (a) Songs vs. Podcasts; (b) Songs vs. Advertisements; and (c) Podcasts vs. Advertisements. Afterward, we created a multi-class classification model using an ensemble technique to differentiate between Songs, Podcasts, and Advertisements in both Hindi and English audio datasets. We have assessed the classification model by employing metrics such as precision, recall, and F1 Score. To conclude, we assessed various classification models and recommended a multi-step strategy for the effective classification of Songs, Podcasts, and Advertisements for Both Hindi and English Audio data. Our models demonstrate outstanding performance in multilingual settings and exhibit enhanced generalization abilities, as confirmed by the analysis of the results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harmonizing Data: Multilingual, Multistep Deep Learning Approach for Classifying Audio Content into Songs, Podcasts (Talk), and Advertisement

  • Chandan Verma,
  • Naveen Saini,
  • Saurabh Shukla

摘要

In the rapidly evolving field of audio analysis, distinguishing between diverse audio formats like Podcasts (Talk), Advertisements and Songs poses unique challenges due to their varying content, structure, and features. Addressing this issue, our paper, introduces a novel approach to classifying these distinct types of audio content through advanced deep learning techniques. To precisely understand the core attributes of each audio format, we deployed a thorough extraction of features through two methods: the first being Mel-Frequency Cepstral Coefficients (MFCCs), and the second, the application of decibel scaling to the Mel spectrogram feature, which transform audio information into spectrograms. Utilizing these spectrograms as a foundation, our study delves into the training of various Convolutional Neural Network (CNN) frameworks, including VGGNet, DenseNet, and MobileNet. We have collected Audio data from various open-source platforms, including Google Podcasts, YouTube, among others. Initially, we developed binary classification models to categorize Hindi and English audio content into three distinct pairs: (a) Songs vs. Podcasts; (b) Songs vs. Advertisements; and (c) Podcasts vs. Advertisements. Afterward, we created a multi-class classification model using an ensemble technique to differentiate between Songs, Podcasts, and Advertisements in both Hindi and English audio datasets. We have assessed the classification model by employing metrics such as precision, recall, and F1 Score. To conclude, we assessed various classification models and recommended a multi-step strategy for the effective classification of Songs, Podcasts, and Advertisements for Both Hindi and English Audio data. Our models demonstrate outstanding performance in multilingual settings and exhibit enhanced generalization abilities, as confirmed by the analysis of the results.