Parkinson’s disease (PD) is a neurological condition that produces several speech deficits, typically known as hypokinetic dysarthria, affecting the production of different phonemes and resulting in an impaired speech communication. This work presents a detailed investigation based on the wav2vec 2.0 foundational model specifically tuned to perform the automatic discrimination between PD and healthy control (HC) subjects. The investigation showed that, instead of considering the complete wav2vec 2.0 architecture with 12 layers, the five layer is enough to find a model suitable to obtain good classification accuracies. Besides, this work presents a framework where frame-wise classification results are considered, enabling a detailed analysis regarding which phonemes and phonological classes are more accurate for performing the classification. All experiments are evaluated in an external and independent test set, therefore given the good results found in this work, which motivates us to continue working in this direction. For future work, we plan to modify the method to perform the time-stamp labeling to model co-articulation information in speech produced by PD patients.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the Use of a Foundation Acoustic Model to Identify Highly Relevant Phonetic Information of Parkinson’s Speech

  • D. Escobar-Grisales,
  • C. D. Ríos-Urrego,
  • J. R. Orozco-Arroyave

摘要

Parkinson’s disease (PD) is a neurological condition that produces several speech deficits, typically known as hypokinetic dysarthria, affecting the production of different phonemes and resulting in an impaired speech communication. This work presents a detailed investigation based on the wav2vec 2.0 foundational model specifically tuned to perform the automatic discrimination between PD and healthy control (HC) subjects. The investigation showed that, instead of considering the complete wav2vec 2.0 architecture with 12 layers, the five layer is enough to find a model suitable to obtain good classification accuracies. Besides, this work presents a framework where frame-wise classification results are considered, enabling a detailed analysis regarding which phonemes and phonological classes are more accurate for performing the classification. All experiments are evaluated in an external and independent test set, therefore given the good results found in this work, which motivates us to continue working in this direction. For future work, we plan to modify the method to perform the time-stamp labeling to model co-articulation information in speech produced by PD patients.