Alzheimer’s Disease (AD) is the major cause of dementia worldwide and is characterized by cognitive decline that affects speech and language. Conventional techniques for detection of AD such as neuroimaging and cerebrospinal fluid analysis are invasive and costly. Our approach uses a state-of-the-art hybrid of deep learning and classical machine learning-based framework for the early detection of AD, by leveraging the ADReSSo 2021 Challenge dataset, which consists of audio recordings from both AD affected and cognitively normal (CN) individuals. This approach is scalable, cost-efficient, and non-invasive. The features were extracted from the audio files by using Wav2Vec 2.0 in a sliding window approach and Principal Component Analysis (PCA) was applied for dimensionality reduction of the feature vector and then passed to models like XGBoost (reaching an accuracy of 84%), Random Forest, MLP Classifier and Gradient Boosting. The text model uses OpenAI’s Whisper model for generating transcripts of the audio files, and these transcripts are preprocessed using PCA and n-gram based TF-IDF Vectorization, further enriched with lexical diversity features such as MATTR and MTLD, to capture subtle linguistic impairments. These models were able to achieve significant results across both text and audio pipeline. This approach provides a scalable, cost-efficient, and non-invasive solution for early Alzheimer’s detection, achieving up to 86% accuracy through the multimodal integration of audio and linguistic features.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Framework for Early Detection of Alzheimer’s Disease Using Acoustic and Linguistic Features

  • K. N. C. Vardhan,
  • R. Jayashree

摘要

Alzheimer’s Disease (AD) is the major cause of dementia worldwide and is characterized by cognitive decline that affects speech and language. Conventional techniques for detection of AD such as neuroimaging and cerebrospinal fluid analysis are invasive and costly. Our approach uses a state-of-the-art hybrid of deep learning and classical machine learning-based framework for the early detection of AD, by leveraging the ADReSSo 2021 Challenge dataset, which consists of audio recordings from both AD affected and cognitively normal (CN) individuals. This approach is scalable, cost-efficient, and non-invasive. The features were extracted from the audio files by using Wav2Vec 2.0 in a sliding window approach and Principal Component Analysis (PCA) was applied for dimensionality reduction of the feature vector and then passed to models like XGBoost (reaching an accuracy of 84%), Random Forest, MLP Classifier and Gradient Boosting. The text model uses OpenAI’s Whisper model for generating transcripts of the audio files, and these transcripts are preprocessed using PCA and n-gram based TF-IDF Vectorization, further enriched with lexical diversity features such as MATTR and MTLD, to capture subtle linguistic impairments. These models were able to achieve significant results across both text and audio pipeline. This approach provides a scalable, cost-efficient, and non-invasive solution for early Alzheimer’s detection, achieving up to 86% accuracy through the multimodal integration of audio and linguistic features.