<p>The CAZy database organizes carbohydrate-active enzymes into more than 800 families with distinct biochemical roles, making automatic family assignment from sequence a high-value task in metagenomics and enzyme engineering. Although recent deep learning studies have reported strong classification performance, we identify a critical evaluation bias that has not been systematically quantified in this setting: commonly used random train/test splits allow homologous sequences to appear in both training and evaluation sets, thereby inflating apparent performance by 5.9–12.0%. We refer to this effect as homology leakage and quantify its magnitude across five representative classification approaches. To enable more rigorous evaluation, we construct a homology-aware benchmark comprising 30,000 sequences, 60 families, and 6 classes, split using MMseqs2 clustering at 20% sequence identity. Under this protocol, the method most strongly affected by leakage is homology-based inference itself (<InlineEquation ID="IEq1"><EquationSource Format="TEX">\(\Delta \)</EquationSource></InlineEquation>F1 <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(= +0.076\)</EquationSource></InlineEquation>), whose performance would be reported as 0.894 under a random split but drops to 0.818 under homology-aware evaluation. Against this benchmark, our proposed model, ESM-2+Mean+MTL (two-stage selective fine-tuning with joint family/class supervision), achieves a family Macro-F1 of 0.805&#xa0;±&#xa0;0.004, narrowing the gap to homology-based inference while providing calibrated uncertainty estimates (ECE <InlineEquation ID="IEq4"><EquationSource Format="TEX">\(= 0.028\)</EquationSource></InlineEquation>) and compatibility with 8&#xa0;GB consumer GPU hardware. These findings support homology-aware evaluation as a more reliable standard for protein sequence classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Homology aware benchmarking of deep learning models for CAZy enzyme family classification

  • Ahmet Haşim Yurttakal,
  • Hasan Erbay

摘要

The CAZy database organizes carbohydrate-active enzymes into more than 800 families with distinct biochemical roles, making automatic family assignment from sequence a high-value task in metagenomics and enzyme engineering. Although recent deep learning studies have reported strong classification performance, we identify a critical evaluation bias that has not been systematically quantified in this setting: commonly used random train/test splits allow homologous sequences to appear in both training and evaluation sets, thereby inflating apparent performance by 5.9–12.0%. We refer to this effect as homology leakage and quantify its magnitude across five representative classification approaches. To enable more rigorous evaluation, we construct a homology-aware benchmark comprising 30,000 sequences, 60 families, and 6 classes, split using MMseqs2 clustering at 20% sequence identity. Under this protocol, the method most strongly affected by leakage is homology-based inference itself (\(\Delta \)F1 \(= +0.076\)), whose performance would be reported as 0.894 under a random split but drops to 0.818 under homology-aware evaluation. Against this benchmark, our proposed model, ESM-2+Mean+MTL (two-stage selective fine-tuning with joint family/class supervision), achieves a family Macro-F1 of 0.805 ± 0.004, narrowing the gap to homology-based inference while providing calibrated uncertainty estimates (ECE \(= 0.028\)) and compatibility with 8 GB consumer GPU hardware. These findings support homology-aware evaluation as a more reliable standard for protein sequence classification.