Long non-coding RNAs (lncRNAs) are increasingly recognized as crucial regulatory molecules in numerous biological processes, yet their identification and functional annotation remain challenging. This paper presents a novel machine-learning pipeline, leveraging natural language processing (NLP) concepts, for distinguishing lncRNA sequences from protein-coding transcripts. We use a k-mer bag-of-words approach to transform RNA sequences into numerical vectors and train a Support Vector Machine (SVM) classifier to predict lncRNA status. Our proposed framework demonstrates high accuracy in classifying validated lncRNA and protein-coding transcripts, offering a promising approach for large-scale lncRNA discovery in diverse organisms. We discuss methodological innovations, statistical insights, and potential biological applications, with an emphasis on how data science drives innovation in genomics research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging NLP-Inspired Support Vector Machine for lncRNA Identification

  • Andrea Bonomo,
  • Elena Ballante,
  • Domenico Giosa,
  • Silvia Figini

摘要

Long non-coding RNAs (lncRNAs) are increasingly recognized as crucial regulatory molecules in numerous biological processes, yet their identification and functional annotation remain challenging. This paper presents a novel machine-learning pipeline, leveraging natural language processing (NLP) concepts, for distinguishing lncRNA sequences from protein-coding transcripts. We use a k-mer bag-of-words approach to transform RNA sequences into numerical vectors and train a Support Vector Machine (SVM) classifier to predict lncRNA status. Our proposed framework demonstrates high accuracy in classifying validated lncRNA and protein-coding transcripts, offering a promising approach for large-scale lncRNA discovery in diverse organisms. We discuss methodological innovations, statistical insights, and potential biological applications, with an emphasis on how data science drives innovation in genomics research.