Life is composed of sequences, but due to the complexity of biological sequences, clustering algorithms have been introduced for the analysis and processing of biological sequence data. However, in tasks involving synthetic DNA sequences, such as DNA data storage, under high-error-rate sequencing techniques like nanopore sequencing, the accuracy of clustering and the reliability of reconstruction remain significant challenges. Therefore, this paper proposes a hash sketches fuzzy clustering (HSFC) method for reliable DNA storage data reconstruction. HSFC employs locality sensitive hashing to map DNA sequences as hash sketches with drifts and designs fuzzy matching mechanisms that tolerate more sequence errors, thereby mitigating the impact of errors on clustering results. Experimental results show that HSFC improves the clustering accuracy of DNA sequences by 6% to 17% compared to state-of-the-art DNA clustering methods. Moreover, HSFC achieves sequence recovery and reconstruction rates of 99% at a simulation error rate of 10%. In conclusion, HSFC enhances the accuracy of DNA sequence clustering in high error rate environments, thus facilitating high quality data reconstruction and ensuring the integrity and reliability of DNA storage read data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DNA Sequence Clustering in High Error Rates via Hash Sketches Fuzzy Clustering for Efficient Stored Data Reconstruction

  • Qi Shao,
  • Yanfen Zheng,
  • Ben Cao,
  • Zhenlu Liu,
  • Bin Wang,
  • Shihua Zhou,
  • Pan Zheng

摘要

Life is composed of sequences, but due to the complexity of biological sequences, clustering algorithms have been introduced for the analysis and processing of biological sequence data. However, in tasks involving synthetic DNA sequences, such as DNA data storage, under high-error-rate sequencing techniques like nanopore sequencing, the accuracy of clustering and the reliability of reconstruction remain significant challenges. Therefore, this paper proposes a hash sketches fuzzy clustering (HSFC) method for reliable DNA storage data reconstruction. HSFC employs locality sensitive hashing to map DNA sequences as hash sketches with drifts and designs fuzzy matching mechanisms that tolerate more sequence errors, thereby mitigating the impact of errors on clustering results. Experimental results show that HSFC improves the clustering accuracy of DNA sequences by 6% to 17% compared to state-of-the-art DNA clustering methods. Moreover, HSFC achieves sequence recovery and reconstruction rates of 99% at a simulation error rate of 10%. In conclusion, HSFC enhances the accuracy of DNA sequence clustering in high error rate environments, thus facilitating high quality data reconstruction and ensuring the integrity and reliability of DNA storage read data.