A Speech Enhancement Method Based on Training Lifetime Knowledge Distillation
摘要
Single-channel speech enhancement (SE) using deep neural networks has achieved impressive results in recent years. However, the large size and high computational cost of state-of-the-art models pose challenges for deployment in resource-constrained environments. To address this issue, we propose a novel knowledge distillation framework named Training Lifetime Knowledge Distillation (TLKD). Unlike traditional methods that distill knowledge from a single teacher snapshot, TLKD leverages multiple intermediate teacher models saved during training and employs a self-attention-based feature integration strategy to effectively transfer dynamic knowledge to a lightweight student model. Experimental results on the VoiceBank-DEMAND dataset demonstrate that the proposed method significantly improves the performance of the student model across various speech quality metrics while maintaining low computational complexity. Additionally, ablation studies show that an appropriate number of intermediate teacher models yields optimal results, highlighting the trade-off between performance and resource consumption. The proposed TLKD method offers a practical and effective solution for deploying high-performance speech enhancement models on low-power devices.