Automatic voice pathology detection and classification using Hamilton product RNN
摘要
Automated detection of pathologic voice with machine learning approaches is a recent research field in the health domain since it needs a laborious, rigorous, and time-consuming examination. This problem can be solved by automatic detection of voice disorders using deep learning. In this study, a powerful feature extraction voice disorder estimation based on the deep learning model is developed. First, the glottal closure instant, the glottal opening instant, the open quotient, the fundamental frequency, and the cepstral features parameters based on the multi-scale product analysis (MPA) of the speech signal are detected. Also, the logarithmic amplitude and phase variation of MPA, considered non-parametric and non-conventional features issued from amplitude and phase in the spectral domain, are measured. These parameter measurements from the MPA of the speech signal are used as features for the Hamilton product Recurrent Neural Network (Ham-RNN) classifier. The classifier considers both the external and internal dependencies and allows to code the dependencies by composing the multi-dimensional features as single entities as well as by determining the correlations between the elements by the recurrent operation. In this study, experiments are performed to analyze the impact of age, gender and regional accents on the model performance, based on the MEEI and SVD datasets with a large panel of voice pathologies. Furthermore, an ablation analysis of the acoustic descriptors allowed to select the most relevant features, followed by a comparison between the model performance on the original data, with and without data augmentation via SMOTE, as well as with the addition of an external database used exclusively for the test phase. The results show the potential of using the phase component and the MPA for automatically detecting pathological voices.