As humanity delves deeper into DNA-related research, the efficient storage of DNA sequence data has emerged as a crucial challenge. Traditional lossless compression methods, such as RLE and LZ algorithms, fail to capture the long-distance repetitiveness and uniform distribution inherent in DNA sequences. To mitigate this issue, we propose DNA-PRIME, a novel lossless DNA compression framework. It employs binary encoding to elevate the probability of recurring patterns. Subsequently, we segment the DNA sequence into blocks with internal similarities using a content-based chunking algorithm. Inspired by SimHash, we further devise a Weight Hash method integrated with a feature weight dictionary, allowing the primary features within a block to be expressed quantifiable. Moreover, we employ feature fusion techniques to accentuate the principal features within blocks and minimize comparison times. Building upon this, we incorporate tree-pruning strategies to enhance overall efficiency. By leveraging the characteristics of the data, we ultimately employ advanced compression methodologies to achieve superior compression rates. Finally, we conduct experiments on real workloads. The results indicate that our framework can cope with the unique challenges of DNA sequence compression to a certain extent, substantially improving both compression efficiency and fidelity to sequence characteristics. Our findings highlight the potential of our approach to enhance DNA data management across various computational biology applications significantly.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DNA-PRIME: Advanced DNA Sequence Compression Through Enhanced Feature Fusion and Weight Hashing

  • Xin Liu,
  • Zhenxi Tian,
  • Wenlong Tian,
  • Zhiyong Xu

摘要

As humanity delves deeper into DNA-related research, the efficient storage of DNA sequence data has emerged as a crucial challenge. Traditional lossless compression methods, such as RLE and LZ algorithms, fail to capture the long-distance repetitiveness and uniform distribution inherent in DNA sequences. To mitigate this issue, we propose DNA-PRIME, a novel lossless DNA compression framework. It employs binary encoding to elevate the probability of recurring patterns. Subsequently, we segment the DNA sequence into blocks with internal similarities using a content-based chunking algorithm. Inspired by SimHash, we further devise a Weight Hash method integrated with a feature weight dictionary, allowing the primary features within a block to be expressed quantifiable. Moreover, we employ feature fusion techniques to accentuate the principal features within blocks and minimize comparison times. Building upon this, we incorporate tree-pruning strategies to enhance overall efficiency. By leveraging the characteristics of the data, we ultimately employ advanced compression methodologies to achieve superior compression rates. Finally, we conduct experiments on real workloads. The results indicate that our framework can cope with the unique challenges of DNA sequence compression to a certain extent, substantially improving both compression efficiency and fidelity to sequence characteristics. Our findings highlight the potential of our approach to enhance DNA data management across various computational biology applications significantly.