Audio-Visual Wake-up Word Spotting Under Noisy and Multi-person Scenarios
摘要
The existing audio-visual wake-up word spotting (AVWWS) methods assume that the audio signal has been aligned with the lip movement video signal of a specific speaker in noisy environments, and are mainly applicable for scenarios with only a single speaker. However, in complex scenarios, there may be multiple people showing up in the video facing the camera simultaneously, and more than one person may be speaking at the same time. Wake-up word spotting in noisy and multi-person scenarios remains relatively under-explored. In this paper, we first propose a Wake-up Word Active Speaker Detection Model (WWASD) to recognize the face that is speaking the wake-up word. Based on the model, we propose two approaches, namely Two-stage detection and Three-stage detection, for audio-visual wake-up word spotting in noisy and multi-person scenarios. We compare the approaches from the perspectives of performance and computational complexity on MISP2021-AVWWS corpus. The best Two-stage detection approach, which contains WWASD and audio-visual wake-up word spotting model, achieves comparable performance against the systems with oracle visual speaker bounding boxes. Three-stage detection, which adds an audio-based single-modality wake-up word model as a front end greatly reduces the computational cost.