<p>In this paper, we address the problem of the calculation of the cepstral peak prominence (CPP) from noisy recordings for disordered voices analysis. The CPP is an effective acoustic measure for disordered voices assessment. It provides a measure of the regularity of the speech signal spectrum. Low perturbed speech signals have more regular spectrum than highly perturbed speech signals and therefore greater CPP values. In the studies devoted to the accuracy and reliability of acoustic measures, the acoustic cues are extracted from clean speech signals. In many situations such as in telehealth, speech signals used for the acoustic analysis of the voice quality are recorded in noisy environment. Conventional filtering fails to improve the accuracy and reliability of the acoustic cues in presence of noise because the signal and noise frequency spectra significantly overlap. Filtering the noisy speech signal removes not only noise but also the components related to the process of voice and speech production impacting negatively the correlation between the acoustic measures and the scores of perceptual rating. In the presence of background noise, the validity and reliability of the acoustic measures for voice quality assessment might be negatively impacted. The objective of this paper is to propose a smoothing method named ensemble compressive sensing (CS) to improve the performance of the CPP in terms of correlation with the degree of perceived hoarseness. The CS is a new method for simultaneous acquisition and compression of signals with sparse representation. The main challenge is to reduce the effect of background noise without altering the vocal noise caused by the dysfunction of the vocal folds. We show that the CS-based smoothing method performs better than the reference methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cepstral analysis of noisy speech for disordered voices assessment

  • Abdellah Kacha,
  • Francis Grenez

摘要

In this paper, we address the problem of the calculation of the cepstral peak prominence (CPP) from noisy recordings for disordered voices analysis. The CPP is an effective acoustic measure for disordered voices assessment. It provides a measure of the regularity of the speech signal spectrum. Low perturbed speech signals have more regular spectrum than highly perturbed speech signals and therefore greater CPP values. In the studies devoted to the accuracy and reliability of acoustic measures, the acoustic cues are extracted from clean speech signals. In many situations such as in telehealth, speech signals used for the acoustic analysis of the voice quality are recorded in noisy environment. Conventional filtering fails to improve the accuracy and reliability of the acoustic cues in presence of noise because the signal and noise frequency spectra significantly overlap. Filtering the noisy speech signal removes not only noise but also the components related to the process of voice and speech production impacting negatively the correlation between the acoustic measures and the scores of perceptual rating. In the presence of background noise, the validity and reliability of the acoustic measures for voice quality assessment might be negatively impacted. The objective of this paper is to propose a smoothing method named ensemble compressive sensing (CS) to improve the performance of the CPP in terms of correlation with the degree of perceived hoarseness. The CS is a new method for simultaneous acquisition and compression of signals with sparse representation. The main challenge is to reduce the effect of background noise without altering the vocal noise caused by the dysfunction of the vocal folds. We show that the CS-based smoothing method performs better than the reference methods.