<p>The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH<sub>opt</sub>) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH<sub>opt</sub> will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence–function relationships. Here we proposed and evaluated various machine learning methods for predicting pH<sub>opt</sub>, conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pH<sub>opt</sub>. We present EpHod, the best-performing model, to predict pH<sub>opt</sub>, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH<sub>opt</sub>, including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH<sub>opt</sub> prediction and will potentially speed up the development of enzyme technologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning prediction of enzyme optimum pH

  • Japheth E. Gado,
  • Matthew Knotts,
  • Ada Y. Shaw,
  • Debora Marks,
  • Nicholas P. Gauthier,
  • Chris Sander,
  • Gregg T. Beckham

摘要

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pHopt) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pHopt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence–function relationships. Here we proposed and evaluated various machine learning methods for predicting pHopt, conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pHopt, including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pHopt prediction and will potentially speed up the development of enzyme technologies.