DNA has emerged as a promising new storage medium due to its high storage capacity, low maintenance cost, and remarkable longevity. However, limitations of current synthesis and sequencing technologies, particularly short lengths of DNA sequences and random access through primers, poses significant challenges for the efficient storage and retrieval of large-scale datasets. In this work, we propose a novel database system termed DNA-HiCapD for DNA-based data storage, which utilizes a hierarchical tree structure with multiplex primers to manage massive data objects, and introduce a Bloom filter in the primer search process to enable efficient and error-tolerant data retrieval. In addition, we design a multiplex primers generation method that adapt to data structures, thereby improving the address space of the storage system. We argue through theoretical proofs to demonstrate that DNA-HiCapD achieves high storage capacity and low data retrieval complexity. Experimental validation shows that the proposed scheme can store millions of files across several datasets and achieves linear time complexity for data retrieval compared to traditional methods. These advantages highlight the potential application of DNA-HiCapD in large-scale DNA-based data storage.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A High Capacity and Efficient Retrieval Database System for DNA-Based Data Storage

  • Zixiao Zhang,
  • Zuqi Liu,
  • Fei Xu

摘要

DNA has emerged as a promising new storage medium due to its high storage capacity, low maintenance cost, and remarkable longevity. However, limitations of current synthesis and sequencing technologies, particularly short lengths of DNA sequences and random access through primers, poses significant challenges for the efficient storage and retrieval of large-scale datasets. In this work, we propose a novel database system termed DNA-HiCapD for DNA-based data storage, which utilizes a hierarchical tree structure with multiplex primers to manage massive data objects, and introduce a Bloom filter in the primer search process to enable efficient and error-tolerant data retrieval. In addition, we design a multiplex primers generation method that adapt to data structures, thereby improving the address space of the storage system. We argue through theoretical proofs to demonstrate that DNA-HiCapD achieves high storage capacity and low data retrieval complexity. Experimental validation shows that the proposed scheme can store millions of files across several datasets and achieves linear time complexity for data retrieval compared to traditional methods. These advantages highlight the potential application of DNA-HiCapD in large-scale DNA-based data storage.