Construction and Experimental Verification of Automatic Classification Process Based on K-Mer Frequency Statistics
摘要
In bioinformatics, k-mer frequency statistics are an important tool for analyzing biological sequences, widely used for sequence alignment, duplicate detection, correction, species identification, and motif discovery. However, when dealing with large-scale data, existing k-mer frequency statistics tools have significant bottlenecks in the use of memory and disk space. Therefore, this article proposes a new k-mer frequency statistical method - KCOSS, aimed at optimizing storage and computing efficiency. When processing sequences with a length not exceeding 14, KCOSS uses static hash tables and composite bijective functions, effectively reducing storage requirements. For sequences longer than 14, KCOSS combines Bloom filters and two-level hash tables (including static hash tables and dynamic cuckoo hash tables) to improve processing speed. In addition, to further reduce memory and disk usage, KCOSS optimizes continuous k-mers by only storing newly emerging bases. This article also implements a lockless thread pool, a lockless segmented Bloom filter, and a lockless compact hash table to reduce competition for shared memory and improve parallel processing performance. The experimental results show that when processing human genome data, KCOSS is 22.91% to 169.90% faster than KMC3 under 24 threads, and 527.62% to 806.43% faster than Jellyfish 2; Under 48 threads, the speed improvement ranges from 77.27% to 170.58% and 529.60% to 675.84%, respectively. In terms of memory consumption, KCOSS is comparable to KMC3, only 16.67% of Jellyfish 2; In terms of hard disk storage, KCOSS requires only 12.40% to 17.84% of Jellyfish 2's space and 15.73% to 30.69% of KMC3's.