Short-term harmonic peaks (STHPs) have been used in a harmonic frequency-based multiple observation likelihood ratio test (Hmfreq-MOLRT) VAD successfully. Through the characteristics of spectral harmonicity, the method boosts the likelihood ratio (LR) scores for voiced frames under low signal-to-noise ratio (SNR) conditions so that the total score of its decision function is high enough to promote the detection of weaker speech frames. However, when the STHPs are seriously disturbed by various noises, the total score is not improved by the method sufficiently and the VAD’s performance starts to fall. In this paper, we present a novel VAD algorithm to solve the problem. In the algorithm, we design a new geometric mean (GM) of likelihood ratios (LRs) located at long-term harmonic peaks (LTHPs) and realize an integration of STHPs and LTHPs using a two-level discriminative weight training framework. Experimental results show the performance of the proposed method has significant improvement on speech/non-speech detection accuracy in comparison with the Hmfreq-MOLRT algorithm. Especially in extremely adverse noisy conditions, such as −10 dB and −5 dB, the improvement is more obvious.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integration of Short-Term and Long-Term Harmonic Peaks in a Two-Level Discriminative Weight Training Framework for Voice Activity Detection

  • YingWei Tan

摘要

Short-term harmonic peaks (STHPs) have been used in a harmonic frequency-based multiple observation likelihood ratio test (Hmfreq-MOLRT) VAD successfully. Through the characteristics of spectral harmonicity, the method boosts the likelihood ratio (LR) scores for voiced frames under low signal-to-noise ratio (SNR) conditions so that the total score of its decision function is high enough to promote the detection of weaker speech frames. However, when the STHPs are seriously disturbed by various noises, the total score is not improved by the method sufficiently and the VAD’s performance starts to fall. In this paper, we present a novel VAD algorithm to solve the problem. In the algorithm, we design a new geometric mean (GM) of likelihood ratios (LRs) located at long-term harmonic peaks (LTHPs) and realize an integration of STHPs and LTHPs using a two-level discriminative weight training framework. Experimental results show the performance of the proposed method has significant improvement on speech/non-speech detection accuracy in comparison with the Hmfreq-MOLRT algorithm. Especially in extremely adverse noisy conditions, such as −10 dB and −5 dB, the improvement is more obvious.