Robust sound-based bird classification using multiple features and random forest classifier
摘要
The research on bird classification from sound gains momentum in ornithology. This sound-based bird classification uses perceptual features with a filter bank in BARK, MEL and Equivalent rectangular bandwidth (ERB) frequency scales and log-energy in Time–frequency (T–F) unit features by taking Cochleagram on the responses of the Gammatone filter bank in BARK, MEL and ERB frequency scales. 80% of these features are given to the modelling technique for developing bootstrap aggregation for an ensemble of decision trees. 20% of the features are given to the model for classifying the bird by predicting the class for each feature vector. The performance of the system is assessed using recognition accuracy as a metric. Decision-level fusion of perceptual features with filters in different frequency scales has provided a maximum accuracy of 93%, besides delivering 100% accuracy for some bird species. Decision-level fusion of Cochleagram features with Gammatone filters in different frequency scales has yielded a maximum accuracy of 99%. This work has considered 42 bird species spanning diverse regions across the world. Ornithologists use this automated system to determine the ecosystem's health.