DNA sequences are the genetic blueprint of an organism and are made up of macromolecular nucleotides. Genes are classified according to their function, structure, and regulation, which aids in identifying genes involved in both normal cellular functioning and diseases such as cancer. Normal genes regulate cell and tissue function, which has an impact on general health. Unique dinucleotide patterns within gene sequences encode a variety of physiological roles. Understanding these patterns is critical to understanding the genetic basis of disease and health. This study looks at dinucleotide patterns in DNA sequences for both normal and cancer-related genes. We determine the normalized probability of nucleotide pairs generating dinucleotides. This attribute implies lowering the feature complexity of DNA sequences to reduce the computational cost associated with gene analysis. Our findings show that by using proper feature selection and normalization procedures, a logistic regression classifier efficiently differentiates between cancerous and non-cancerous genomic data. The suggested method was successfully tested on relevant research case studies. Overall, the suggested normalized dinucleotide patterns for the feature reduction and logistic regression classifier combination effectively discriminate between cancerous and non-cancerous genes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Machine Learning Model to Differentiating Normal and Cancer Genes Using Dinucleotide Analysis

  • Vijay Kalal,
  • Brajesh Kumar Jha

摘要

DNA sequences are the genetic blueprint of an organism and are made up of macromolecular nucleotides. Genes are classified according to their function, structure, and regulation, which aids in identifying genes involved in both normal cellular functioning and diseases such as cancer. Normal genes regulate cell and tissue function, which has an impact on general health. Unique dinucleotide patterns within gene sequences encode a variety of physiological roles. Understanding these patterns is critical to understanding the genetic basis of disease and health. This study looks at dinucleotide patterns in DNA sequences for both normal and cancer-related genes. We determine the normalized probability of nucleotide pairs generating dinucleotides. This attribute implies lowering the feature complexity of DNA sequences to reduce the computational cost associated with gene analysis. Our findings show that by using proper feature selection and normalization procedures, a logistic regression classifier efficiently differentiates between cancerous and non-cancerous genomic data. The suggested method was successfully tested on relevant research case studies. Overall, the suggested normalized dinucleotide patterns for the feature reduction and logistic regression classifier combination effectively discriminate between cancerous and non-cancerous genes.