The rapid evolution of Next-Generation Sequencing technologies has transformed genomics, generating vast amounts of data that require advanced analysis methods. This paper presents a Deep Learning model developed to predict the pathogenicity of genetic variants, a crucial step towards personalized medicine. Our model is trained on a dataset generated with the analysis of a sample of Next-Generation Sequencing outputs, incorporating a mix of clearly labeled and less certain genetic variants. By adopting a semi-supervised learning approach, the model effectively leverages both soft and hard-labelled data. The core of our methodology is the Feature Tokenizer Transformer architecture, which processes numerical and categorical genomic data. Our preprocessing strategy includes several steps to ensure data quality, such as imputation, scaling, and encoding. Our results demonstrate the model’s exceptional accuracy, especially in identifying hard-labelled variants. The significance of the model’s outputs on the soft labels is also discussed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transformer-Enhanced Pathogenicity Prediction with Soft Labels in a Semi-supervised Setup

  • Pablo Enrique Guillem,
  • Marco Zurdo-Tabernero,
  • Liliana Durón Figueroa,
  • Ángel Canal-Alonso,
  • Guillermo Hernández,
  • Angélica González Arrieta,
  • Fernando de la Prieta

摘要

The rapid evolution of Next-Generation Sequencing technologies has transformed genomics, generating vast amounts of data that require advanced analysis methods. This paper presents a Deep Learning model developed to predict the pathogenicity of genetic variants, a crucial step towards personalized medicine. Our model is trained on a dataset generated with the analysis of a sample of Next-Generation Sequencing outputs, incorporating a mix of clearly labeled and less certain genetic variants. By adopting a semi-supervised learning approach, the model effectively leverages both soft and hard-labelled data. The core of our methodology is the Feature Tokenizer Transformer architecture, which processes numerical and categorical genomic data. Our preprocessing strategy includes several steps to ensure data quality, such as imputation, scaling, and encoding. Our results demonstrate the model’s exceptional accuracy, especially in identifying hard-labelled variants. The significance of the model’s outputs on the soft labels is also discussed.