Introduction <p>Parkinson’s disease (PD) is a slow-progressing neurological disorder that usually appears in elderly individuals, although it develops much earlier. While many researchers have focused on detecting PD based on symptoms, only limited research has explored the use of protein sequences for early detection.</p> Methods <p>In this work, we propose a deep sequential learning model, Parkinson’s Protein Hybrid Net (PPHN), which leverages long short-term memory (LSTM) and gated recurrent unit (GRU) to detect PD from protein amino acid sequences. The Parkinson’s and healthy protein sequences were collected from the NCBI and UniProt databases. The relevant features of the amino acid sequences were extracted using three primary descriptors, viz., amino acid composition (AAC), dipeptide composition (DPC), and tripeptide composition (TPC). To capture more comprehensive sequence-level information, composite descriptors (viz., AAC-DPC-TPC, AAC-DPC, AAC-TPC, and DPC-TPC) were further constructed by combining these individual feature sets. Additionally, SHapley Additive exPlanations (SHAP) analysis was incorporated to interpret the model’s predictions and identify the most influential features contributing to classification.</p> Results <p>The proposed method is found to produce promising results for most of the feature descriptors, achieving the highest accuracy of 98.73% with a precision of 0.9890, a recall of 0.9850, and an <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(F_1\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>F</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> score of 0.9900 (for AAC-DPC-TPC feature descriptor compared with other counterpart methods). The results of the paired <i>t</i>-tests and the Wilcoxon signed-rank tests also confirm the statistical significance of the better results obtained by the proposed model versus other compared methods in most cases. Confidence interval (<i>CI</i>) tests performed at 90%, 95%, and 99% confidence levels demonstrate lower error rates and smaller error bounds achieved by the proposed method than those of the other methods for most of the feature descriptors.</p> Conclusion <p>Therefore, the proposed method may serve as an effective computational tool for the early detection of PD from protein sequences.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PPHN: Deep Sequential Learning Architecture for Parkinson’s Disease Prediction using Amino Acid Descriptors from Protein Sequences

  • Bornali Baruah,
  • Ansuman Kumar,
  • Anindya Halder

摘要

Introduction

Parkinson’s disease (PD) is a slow-progressing neurological disorder that usually appears in elderly individuals, although it develops much earlier. While many researchers have focused on detecting PD based on symptoms, only limited research has explored the use of protein sequences for early detection.

Methods

In this work, we propose a deep sequential learning model, Parkinson’s Protein Hybrid Net (PPHN), which leverages long short-term memory (LSTM) and gated recurrent unit (GRU) to detect PD from protein amino acid sequences. The Parkinson’s and healthy protein sequences were collected from the NCBI and UniProt databases. The relevant features of the amino acid sequences were extracted using three primary descriptors, viz., amino acid composition (AAC), dipeptide composition (DPC), and tripeptide composition (TPC). To capture more comprehensive sequence-level information, composite descriptors (viz., AAC-DPC-TPC, AAC-DPC, AAC-TPC, and DPC-TPC) were further constructed by combining these individual feature sets. Additionally, SHapley Additive exPlanations (SHAP) analysis was incorporated to interpret the model’s predictions and identify the most influential features contributing to classification.

Results

The proposed method is found to produce promising results for most of the feature descriptors, achieving the highest accuracy of 98.73% with a precision of 0.9890, a recall of 0.9850, and an \(F_1\) F 1 score of 0.9900 (for AAC-DPC-TPC feature descriptor compared with other counterpart methods). The results of the paired t-tests and the Wilcoxon signed-rank tests also confirm the statistical significance of the better results obtained by the proposed model versus other compared methods in most cases. Confidence interval (CI) tests performed at 90%, 95%, and 99% confidence levels demonstrate lower error rates and smaller error bounds achieved by the proposed method than those of the other methods for most of the feature descriptors.

Conclusion

Therefore, the proposed method may serve as an effective computational tool for the early detection of PD from protein sequences.