Single-channel speech enhancement (SE) using deep neural networks has achieved impressive results in recent years. However, the large size and high computational cost of state-of-the-art models pose challenges for deployment in resource-constrained environments. To address this issue, we propose a novel knowledge distillation framework named Training Lifetime Knowledge Distillation (TLKD). Unlike traditional methods that distill knowledge from a single teacher snapshot, TLKD leverages multiple intermediate teacher models saved during training and employs a self-attention-based feature integration strategy to effectively transfer dynamic knowledge to a lightweight student model. Experimental results on the VoiceBank-DEMAND dataset demonstrate that the proposed method significantly improves the performance of the student model across various speech quality metrics while maintaining low computational complexity. Additionally, ablation studies show that an appropriate number of intermediate teacher models yields optimal results, highlighting the trade-off between performance and resource consumption. The proposed TLKD method offers a practical and effective solution for deploying high-performance speech enhancement models on low-power devices.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Speech Enhancement Method Based on Training Lifetime Knowledge Distillation

  • Tao Zheng,
  • Liejun Wang,
  • Yinfeng Yu

摘要

Single-channel speech enhancement (SE) using deep neural networks has achieved impressive results in recent years. However, the large size and high computational cost of state-of-the-art models pose challenges for deployment in resource-constrained environments. To address this issue, we propose a novel knowledge distillation framework named Training Lifetime Knowledge Distillation (TLKD). Unlike traditional methods that distill knowledge from a single teacher snapshot, TLKD leverages multiple intermediate teacher models saved during training and employs a self-attention-based feature integration strategy to effectively transfer dynamic knowledge to a lightweight student model. Experimental results on the VoiceBank-DEMAND dataset demonstrate that the proposed method significantly improves the performance of the student model across various speech quality metrics while maintaining low computational complexity. Additionally, ablation studies show that an appropriate number of intermediate teacher models yields optimal results, highlighting the trade-off between performance and resource consumption. The proposed TLKD method offers a practical and effective solution for deploying high-performance speech enhancement models on low-power devices.