Pre-training and Adverse Audio Samples for Data-Efficient Wake Word Detection
摘要
This work investigates the impact of pre-training and the use of adverse audio samples on both the data efficiency and performance of end-to-end neural Wake Word Detection systems. Alongside intensive data augmentation, the proposed methodology involves pre-training Keyword Spotting models, followed by fine-tuning to recognize specific wake words by leveraging their foundational capabilities. The study also examines the inclusion of adverse audio samples resembling the target wake word. Experiments evaluate various state-of-the-art architectures to assess the effects of model size, amount of training data, model pre-training, and the incorporation of adverse audio samples on system performance. Results demonstrate that pre-training improves performance, with fine-tuned models consistently outperforming those trained from scratch, especially with limited data. Additionally, training with adverse samples resembling the wake word also enhances results by reducing false acceptance rates. These findings provide valuable insights for developing data-efficient Wake Word Detection systems.