Domain adaptative keyword spotting with multimodal enhancement
摘要
This paper introduces a domain-adaptive visual enhanced keyword spotting framework that overcomes the severe performance degradation of audio-only systems in noisy or distant conditions. The approach employs a visual-acoustic memory architecture, in which visual key memories from observed lip movements query a shared audio memory to generate candidate acoustic representations. These enriched visual embeddings are fused through a reflective attention mechanism that extracts bidirectional alignment signals from a unified attention map. Long-range temporal dynamics are modeled by a continuity-preserving attention network that enforces consistency via future feature prediction and input reconstruction objectives. Evaluations on the Multimodal Information Based Speech Processing challenge corpus show relative improvements of 24.2% on the development set and 14.8% on the test set compared with state-of-the-art baselines, demonstrating marked gains in both recognition accuracy and robustness under adverse conditions.