Impact of Different Time-Spans in Multimodal Representations with Speech and Language Embeddings for Parkinson’s Disease Detection
摘要
Parkinson’s disease (PD) affects millions of people worldwide. Among others, its symptoms negatively impact speech and language production. Therefore, the research community has worked on modeling these modalities (or sources of information) to perform the automatic detection and evaluation of the disease. Although language encodes relevant information regarding the cognitive status of the patients, most works have focused attention on speech production, and few studies have considered language abnormalities. This chapter presents a methodology where speech recordings of PD patients are modeled with speech embeddings resulting from Convolutional Neural Networks (CNNs), and the corresponding transliterations of those recordings are modeled with Bidirectional Encoder Representations from Transformer (BERT) embeddings and CNN architecture adapted to language analysis. First, we evaluated the suitability of these representations to classify PD patients and healthy subjects considering each modality separately, and second, speech and language representations are combined. To perform the combination, speech representations are converted into a static representation following two strategies: statistical functionals and Gaussian Mixture Model Supervectors (GMM-Supervectors). We evaluated the fusion of both information sources on different time-span levels (namely, granularities) to validate whether a finer granularity fusion provides better performance. The results show that speech and language representations are clearly complementary to each other, which motivates us to continue working on different fusion strategies and also on different methods to extract speech and language information.