In view of the imbalance of protein family data and the insufficient accuracy of prediction models, this article adopts the long-short term memory (LSTM) model and uses its unique gating mechanism to capture long-distance dependencies in protein amino acid sequences, thereby improving the accuracy of mutation impact prediction. This article first converts the protein amino acid sequence into a numerical feature vector, including the use of position-specific scoring matrix (PSSM) and derived features such as amino acid physicochemical properties. One-hot encoding is then used to convert the amino acid sequence into a format suitable for LSTM network input. A network structure including a bidirectional LSTM layer is designed to make full use of the contextual information in the sequence and enhance the model’s ability to capture long-distance interactions. The LSTM layer is followed by a fully connected layer to map the extracted features to the output space of protein function prediction. Finally, the Adam optimizer and the cross-entropy loss function are used to train the model to reduce the prediction error. Experiments show that the LSTM model can well predict changes in protein structure and function caused by disease-related mutations, with an accuracy of 94.55%. In addition, the LSTM model can handle the imbalance in protein family data well and can well reflect the advantages of the LSTM model in predicting disease gene mutations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of Machine Learning to Predict Impact of Disease-Related Mutations on Protein Structure and Function

  • Xiang Qi

摘要

In view of the imbalance of protein family data and the insufficient accuracy of prediction models, this article adopts the long-short term memory (LSTM) model and uses its unique gating mechanism to capture long-distance dependencies in protein amino acid sequences, thereby improving the accuracy of mutation impact prediction. This article first converts the protein amino acid sequence into a numerical feature vector, including the use of position-specific scoring matrix (PSSM) and derived features such as amino acid physicochemical properties. One-hot encoding is then used to convert the amino acid sequence into a format suitable for LSTM network input. A network structure including a bidirectional LSTM layer is designed to make full use of the contextual information in the sequence and enhance the model’s ability to capture long-distance interactions. The LSTM layer is followed by a fully connected layer to map the extracted features to the output space of protein function prediction. Finally, the Adam optimizer and the cross-entropy loss function are used to train the model to reduce the prediction error. Experiments show that the LSTM model can well predict changes in protein structure and function caused by disease-related mutations, with an accuracy of 94.55%. In addition, the LSTM model can handle the imbalance in protein family data well and can well reflect the advantages of the LSTM model in predicting disease gene mutations.