<p>It is essential to identify bird species by their calls for ecological monitoring and often involves expert analysis. This paper illustrates the development of a bird sound classification model using the Mel-Frequency Cepstral Coefficients and deep learning architectures, namely Convolutional Neural Network[CNN], Recurrent Neural Network [RNN], Long Short-Term Memory [LSTM], and Bidirectional Long Short-Term Memory [BiLSTM]. In this study, we have collected the dataset from Kaggle and, after preprocessing using noise reduction and amplitude normalization techniques, the MFCC features extracted were used for training all four models to compare their performances. Among these, BiLSTM achieved the highest validation accuracy of 92.68%, compared with CNN, RNN, and LSTM, because this captures better long-term temporal dependencies in birds’ vocalizations. The trained BiLSTM model was deployed in real time on a Raspberry Pi 3, yielding a processing time of 20–35&#xa0;s per clip. This demonstrates that classifiers based on deep learning models can work effectively in the area of field-based biodiversity research using affordable hardware.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep learning approaches for automated bird sound classification using MFCC features

  • R. Shashidhar,
  • Nischal G. Nayaka,
  • S. Chethan,
  • J. Madhura

摘要

It is essential to identify bird species by their calls for ecological monitoring and often involves expert analysis. This paper illustrates the development of a bird sound classification model using the Mel-Frequency Cepstral Coefficients and deep learning architectures, namely Convolutional Neural Network[CNN], Recurrent Neural Network [RNN], Long Short-Term Memory [LSTM], and Bidirectional Long Short-Term Memory [BiLSTM]. In this study, we have collected the dataset from Kaggle and, after preprocessing using noise reduction and amplitude normalization techniques, the MFCC features extracted were used for training all four models to compare their performances. Among these, BiLSTM achieved the highest validation accuracy of 92.68%, compared with CNN, RNN, and LSTM, because this captures better long-term temporal dependencies in birds’ vocalizations. The trained BiLSTM model was deployed in real time on a Raspberry Pi 3, yielding a processing time of 20–35 s per clip. This demonstrates that classifiers based on deep learning models can work effectively in the area of field-based biodiversity research using affordable hardware.