The field of DNA sequencing was discovered with rapid advance technology and as a result, the significant number of genetic data increases requiring robust computational methods. This study delves into utilizing machine learning (ML) models to enhance the classification of DNA sequences. DNA classification is crucial in genomics for species identification, understanding evolutionary relationships, and studying gene functions. We examine machine learning techniques such as nearest neighbors, Gaussian process classification, decision trees, random forests, neural networks, AdaBoost, and support vector machines, highlighting their specific advantages in genomic analysis. Our research focuses on how these models can effectively handle large genomic datasets, improve computational efficiency, and enhance predictive accuracy in various fields, from biomedicine to agriculture and forensics. We critically evaluate each model’s strengths and limitations and provide insights into their practical applications, ultimately guiding future enhancements in DNA classification methodologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing DNA Classification with Machine Learning Models

  • Sultanul Arifeen Hamim,
  • Mohammad Rabiul Islam,
  • Mubasshar U. I. Tamim,
  • Prottoy Prodhan Joy,
  • Soily Ghosh Sneha,
  • Sumaiya Malik

摘要

The field of DNA sequencing was discovered with rapid advance technology and as a result, the significant number of genetic data increases requiring robust computational methods. This study delves into utilizing machine learning (ML) models to enhance the classification of DNA sequences. DNA classification is crucial in genomics for species identification, understanding evolutionary relationships, and studying gene functions. We examine machine learning techniques such as nearest neighbors, Gaussian process classification, decision trees, random forests, neural networks, AdaBoost, and support vector machines, highlighting their specific advantages in genomic analysis. Our research focuses on how these models can effectively handle large genomic datasets, improve computational efficiency, and enhance predictive accuracy in various fields, from biomedicine to agriculture and forensics. We critically evaluate each model’s strengths and limitations and provide insights into their practical applications, ultimately guiding future enhancements in DNA classification methodologies.