<p>The estimation of the Fundamental Frequency (F0) plays a crucial role in various speech processing applications such as speech synthesis, speaker recognition, and health diagnostics. However, accurate F0 extraction becomes highly challenging when speech signals are corrupted by noise or exhibit extreme pitch variations. Traditional methods often fail to maintain robustness in such conditions. In this study, we propose a novel approach leveraging a Convolutional Neural Network (CNN) combined with deep feature loss to generate an electroglottograph (EGG) waveform model, widely regarded as a reliable reference for pitch extraction. The generated EGG signal is later passed to the autocorrelation function to obtain pitch estimates. Our model is extensively evaluated across clean speech and a range of noisy environments such as babble and white noise types with 0–30 dB SNR levels and diverse pitch-varying scenarios such as singing, emotional speech, and pathological conditions. The results demonstrate a significant improvement in F0 estimation accuracy under low signal-to-noise ratios (SNRs) and in cases of substantial pitch variation, outperforming conventional algorithms. The average Gross Pitch Error (GPE) for babble noise is 6% lower than that of other comparative algorithms. Similarly, the Voicing Decision Error (VDE) and F0 Frame Error (FFE) exhibit consistent improvement under babble noise conditions, even at 0&#xa0;dB. Additionally, the GPE score for the soprano singing category is 3% lower, while for emotional speech samples, it is 1% lower. These findings suggest that the proposed model offers a more robust and accurate solution for F0 estimation in real-world speech applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Synthesized Electroglottograph Signals for F0 Estimation using Deep Feature Loss Network

  • Supritha M. Shetty,
  • K. T. Deepak

摘要

The estimation of the Fundamental Frequency (F0) plays a crucial role in various speech processing applications such as speech synthesis, speaker recognition, and health diagnostics. However, accurate F0 extraction becomes highly challenging when speech signals are corrupted by noise or exhibit extreme pitch variations. Traditional methods often fail to maintain robustness in such conditions. In this study, we propose a novel approach leveraging a Convolutional Neural Network (CNN) combined with deep feature loss to generate an electroglottograph (EGG) waveform model, widely regarded as a reliable reference for pitch extraction. The generated EGG signal is later passed to the autocorrelation function to obtain pitch estimates. Our model is extensively evaluated across clean speech and a range of noisy environments such as babble and white noise types with 0–30 dB SNR levels and diverse pitch-varying scenarios such as singing, emotional speech, and pathological conditions. The results demonstrate a significant improvement in F0 estimation accuracy under low signal-to-noise ratios (SNRs) and in cases of substantial pitch variation, outperforming conventional algorithms. The average Gross Pitch Error (GPE) for babble noise is 6% lower than that of other comparative algorithms. Similarly, the Voicing Decision Error (VDE) and F0 Frame Error (FFE) exhibit consistent improvement under babble noise conditions, even at 0 dB. Additionally, the GPE score for the soprano singing category is 3% lower, while for emotional speech samples, it is 1% lower. These findings suggest that the proposed model offers a more robust and accurate solution for F0 estimation in real-world speech applications.