Unsupervised Weak Speech Enhancement Using Periodic Mixing Invariant Training
摘要
Weak speech enhancement technology can improve the clarity and intelligibility of low-intensity speech in noisy environments, reduce the impact of background noise, and improve measurement accuracy. In this paper, we propose a joint time-frequency Cycle-Mixed Invariant Training (TF-Cycle-MixIT) method for unsupervised speech enhancement, specifically designed to address the challenge of enhancing weak speech signals in strong background noise. This approach integrates the MixIT network with the cyclic continuous learning method, overcoming the limitations of traditional speech separation models that heavily depend on pure speech signals as supervisory data. By fusing harmonic features with time-domain information, the model gains a deeper understanding of the intrinsic structure of speech signals, enabling precise denoising in complex acoustic environments. Field experiments conducted in a highway tunnel with strong background noise interference demonstrate the method’s effectiveness in recovering weak speech signals captured by a Distributed Acoustic Sensing system(DAS). The results show that our approach outperforms existing state-of-the-art (SOTA) unsupervised speech enhancement methods on the TIMIT dataset.