Environmental acoustic intelligence through sound event localization and detection: a review
摘要
Sound Event Localization and Detection (SELD) is a critical capability for environmental acoustic intelligence, enabling systems to jointly identify what sounds are active and where they originate. This technology is foundational for applications ranging from smart-city monitoring and autonomous systems to immersive media. Modern SELD research is dominated by deep learning approaches that leverage multi-channel audio, explicit spatial representations, and unified output formats that resolve complex and densely polyphonic scenes. This review provides a comprehensive synthesis of the field, charting its progress from foundational concepts to the state-of-the-art. We systematically analyze key methodological advancements, spatially-informed feature engineering, and the sophisticated data augmentation pipelines that underpin top-performing systems on public benchmarks. This review also highlights emerging opportunities for future research such as distance-aware 3-D SELD and the advancement of data-efficient learning paradigms. By translating these challenges into concrete research directions, this work aims to accelerate the progress of SELD toward robust, field-ready environmental intelligence tools.