<p>Analyzing the similarities among protein sequences is an indispensable task in bioinformatics. The increasing volume of protein sequences stored in databases has prompted a more comprehensive exploration and understanding of these sequences. At the same time, alignment-based methods, require significant computational time for classification. This surge in data availability has led to the development of alignment-free methods capable of efficiently managing and analyzing diverse datasets. In the proposed alignment-free method, the value of hydrophobicity, a physicochemical property is combined with the probabilities of amino acids in the protein sequence to construct random variables. Afterward, the Euclidean distance measure is used to calculate the distance matrix. Next, phylogenetic trees are generated for protein sequences of ND5, Mammalian, Coronavirus Spike proteins, Beta Globin and Human Rhinovirus virus using the NJ clustering method. These trees are then compared with trees derived using alternative methods viz. Fitch-Margoliash, PM technique, FI method, MSA method, MPT method, NJR method, MCG method, ODM method, MOI method, k-mer based method and Machine Learning based method. Moreover, the validation process using SD and RF value involved comparing this current approach to the Clustal Omega method, considering symmetric distance and computational time. The findings depict that the suggested alignment-free method offers accurate species classifications and increased computational efficiency compared to other methods. Therefore, this method delivers a new perspective on protein sequence analysis, potentially providing advantages over traditional methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Random Variable Based Alignment-Free Approach for Protein Sequence Comparison

  • Debrupa Pal,
  • Papri Ghosh,
  • Subhram Das,
  • Bansibadan Maji

摘要

Analyzing the similarities among protein sequences is an indispensable task in bioinformatics. The increasing volume of protein sequences stored in databases has prompted a more comprehensive exploration and understanding of these sequences. At the same time, alignment-based methods, require significant computational time for classification. This surge in data availability has led to the development of alignment-free methods capable of efficiently managing and analyzing diverse datasets. In the proposed alignment-free method, the value of hydrophobicity, a physicochemical property is combined with the probabilities of amino acids in the protein sequence to construct random variables. Afterward, the Euclidean distance measure is used to calculate the distance matrix. Next, phylogenetic trees are generated for protein sequences of ND5, Mammalian, Coronavirus Spike proteins, Beta Globin and Human Rhinovirus virus using the NJ clustering method. These trees are then compared with trees derived using alternative methods viz. Fitch-Margoliash, PM technique, FI method, MSA method, MPT method, NJR method, MCG method, ODM method, MOI method, k-mer based method and Machine Learning based method. Moreover, the validation process using SD and RF value involved comparing this current approach to the Clustal Omega method, considering symmetric distance and computational time. The findings depict that the suggested alignment-free method offers accurate species classifications and increased computational efficiency compared to other methods. Therefore, this method delivers a new perspective on protein sequence analysis, potentially providing advantages over traditional methods.