Abstract <p>The identification of protein homologs in large databases is critical for biological advancements. Traditional methods, such as protein sequence alignment, often miss remote homologs. To address this limitation, we present the Basic Embedding Search Tool (BEST), a fast and sensitive approach that employs protein language models to create sequence embeddings enriched with evolutionary and structural information. Besides, we introduce a segmented distillation pruning technique to accelerate sequence encoding and develop a multi-layer acceleration structure to achieve a 4290.86-fold speedup in swift access and retrieval of dense vectors. Extensive experiments on real datasets demonstrate that BEST increases sensitivity by over 20% compared to prior methods while maintaining precision and recall. It operates 23.41 times faster than traditional tools like PSI-BLAST and 3.92 times faster than Foldseek, while also detecting homologous sequences that conventional methods miss. BEST and its open-access web server (<a href="http://pm2s.cpolar.top/best1/">http://pm2s.cpolar.top/best1/</a>) are poised to significantly aid enzyme mining and advance biological research. The code is publicly available at <a href="https://github.com/SkyTai-W/ProteinMiningEvaluator">https://github.com/SkyTai-W/ProteinMiningEvaluator</a>.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BEST: Basic Embedding Search Tool Enhancing Discovery of Novel Enzyme

  • Yuxuan Wu,
  • Xiao Yi,
  • Yang Tan,
  • Huiqun Yu,
  • Guisheng Fan,
  • Gaowei Zheng

摘要

Abstract

The identification of protein homologs in large databases is critical for biological advancements. Traditional methods, such as protein sequence alignment, often miss remote homologs. To address this limitation, we present the Basic Embedding Search Tool (BEST), a fast and sensitive approach that employs protein language models to create sequence embeddings enriched with evolutionary and structural information. Besides, we introduce a segmented distillation pruning technique to accelerate sequence encoding and develop a multi-layer acceleration structure to achieve a 4290.86-fold speedup in swift access and retrieval of dense vectors. Extensive experiments on real datasets demonstrate that BEST increases sensitivity by over 20% compared to prior methods while maintaining precision and recall. It operates 23.41 times faster than traditional tools like PSI-BLAST and 3.92 times faster than Foldseek, while also detecting homologous sequences that conventional methods miss. BEST and its open-access web server (http://pm2s.cpolar.top/best1/) are poised to significantly aid enzyme mining and advance biological research. The code is publicly available at https://github.com/SkyTai-W/ProteinMiningEvaluator.

Graphical Abstract