SeqAlignXGBoost: Sequence Alignment and Feature Selection for m1A Modification Site Identification
摘要
The rapid advancement of high-throughput sequencing technologies has facilitated extensive research into RNA modifications, particularly N1-methyladenosine (m1A), which plays a critical role in RNA stability, translation, and cellular functions. However, accurately identifying m1A modification sites from large-scale RNA datasets remains challenging. To address this, we propose SeqAlignXGBoost, a computational framework that integrates advanced feature extraction techniques—including k-mer features, Position Specific k-mer Nucleotide Composition (PseKNC), Pseudo-Discrete Nucleotide Composition (PseDNC), and sequence alignment features—with the XGBoost algorithm for efficient m1A site prediction. The model employs incremental feature selection to optimize performance and reduce computational complexity. Experimental results demonstrate that SeqAlignXGBoost outperforms traditional machine learning models, achieving a balanced performance. SeqAlignXGBoost provides a robust and scalable solution for m1A modification site identification, advancing RNA modification research in genomics and bioinformatics.