A High Capacity and Efficient Retrieval Database System for DNA-Based Data Storage
摘要
DNA has emerged as a promising new storage medium due to its high storage capacity, low maintenance cost, and remarkable longevity. However, limitations of current synthesis and sequencing technologies, particularly short lengths of DNA sequences and random access through primers, poses significant challenges for the efficient storage and retrieval of large-scale datasets. In this work, we propose a novel database system termed DNA-HiCapD for DNA-based data storage, which utilizes a hierarchical tree structure with multiplex primers to manage massive data objects, and introduce a Bloom filter in the primer search process to enable efficient and error-tolerant data retrieval. In addition, we design a multiplex primers generation method that adapt to data structures, thereby improving the address space of the storage system. We argue through theoretical proofs to demonstrate that DNA-HiCapD achieves high storage capacity and low data retrieval complexity. Experimental validation shows that the proposed scheme can store millions of files across several datasets and achieves linear time complexity for data retrieval compared to traditional methods. These advantages highlight the potential application of DNA-HiCapD in large-scale DNA-based data storage.