Human Leukocyte antigen (HLA) molecules are the principal defining element in T-cell immunogenicity and play a central role in the recognition of antigen peptides. Accurate estimations of the interactions between antigen peptides and HLA (pHLA) is crucial to facilitate vaccine development and active immunotherapies. Recent advances in pre-trained protein language models (PLMs) have proven their capabilities for protein structure prediction tasks. However, PLMs require high GPU-memory and expensive computation power to train, leading to over-parameterized models. We propose PAMPHLATE, a parameter-efficient and fine-tuned PLM for the prediction of pHLA interactions. It leverages ESM2 and achieves highly accurate pHLA predictions through domain adaptation and reparameterization techniques. We reduce the model parameters by 99-fold with a 2.5x reduction in time over vanilla fine-tuning, while outperforming several state-of-the-art methods. Our model achieves accuracy of 0.88 and MCC of 0.76 on an external dataset. PAMPHLATE secures the highest true positives and lowest false negatives for a neoantigen and HPV vaccine dataset showcasing its superiority over existing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ParAMeter Efficient Fine-Tuned Language Models for Peptide-HLA InTeraction PrEdiction

  • Hamda Alhosani,
  • Raghvendra Mall,
  • Ankita Singh,
  • Filippo Castiglione

摘要

Human Leukocyte antigen (HLA) molecules are the principal defining element in T-cell immunogenicity and play a central role in the recognition of antigen peptides. Accurate estimations of the interactions between antigen peptides and HLA (pHLA) is crucial to facilitate vaccine development and active immunotherapies. Recent advances in pre-trained protein language models (PLMs) have proven their capabilities for protein structure prediction tasks. However, PLMs require high GPU-memory and expensive computation power to train, leading to over-parameterized models. We propose PAMPHLATE, a parameter-efficient and fine-tuned PLM for the prediction of pHLA interactions. It leverages ESM2 and achieves highly accurate pHLA predictions through domain adaptation and reparameterization techniques. We reduce the model parameters by 99-fold with a 2.5x reduction in time over vanilla fine-tuning, while outperforming several state-of-the-art methods. Our model achieves accuracy of 0.88 and MCC of 0.76 on an external dataset. PAMPHLATE secures the highest true positives and lowest false negatives for a neoantigen and HPV vaccine dataset showcasing its superiority over existing methods.