Multi-species Acoustic Recognition with Deep Learning and Cloud Technology for Comprehensive Biodiversity Monitoring
摘要
The decline in fauna populations is a cause for concern, not just for ecosystems but for our own well-being too. It’s not just about the animals themselves; their disappearance disrupts entire ecosystems, affecting plants, food chains, and crucial services like clean water and food security. The problem is that it’s hard for zoologists to identify which animal is most at risk of disappearing forever. This makes it tough to come up with plans to identify them in the manual process. So, our work has focused on implementing a robust framework incorporating API integration, data integrity methods, spectrogram analysis, waveform analysis, and audio data augmentation (time shifting), as well as spectrogram augmentation (Spec Augment) techniques. Leveraging advanced methods such as MFCC integrated with zero crossing rate, which efficiently extracts valuable insights from the acoustic data produced by the fauna, we’ve been using deep neural networks called SpectraFusionNet (GRU + VGGish) to make predictions on fauna acoustic data. Plus, we’ve made it accessible by deploying our models into a Streamlit community cloud, so anyone can use it. Our goal isn’t just to help researchers; it’s to bring everyone together in conversation about conservation. By making real-time acoustic data available to everyone, we’re giving people the power to make a difference in protecting our planet’s biodiversity. It’s about using technology to bridge the gap between innovation and sustainability, and ultimately, it’s about empowering people to be part of the solution.