Advanced chunk-based data deduplication framework for secure data storage in cloud using hybrid heuristic assisted optimal key-based encryption
摘要
In a cloud storage system, the information is stored and secured through several encryption methods. Data redundancy is an important problem that affects data in the cloud and it enlarges large space in storage environments. Storing and managing a massive volume of information is exceedingly difficult. The existing method could not have the ability to remove the redundant data blocks during the authentication process. During the deduplication process, the security and ownership of data are highly affected. In order to identify high levels of redundancy, Content-Defined Chunking (CDC) plays a significant role in the redundant nature of data solutions. Several deduplication approaches have been developed to rectify these issues. However, most of the approaches are affected by security flaws, since they don’t examine with dynamic scenarios in the cloud storage environment. Chunk-based data deduplication approach is developed in this research work to provide high security for data in cloud. It preserves the cloud storage to analyze repeated files during the file upload. The repeated files are eliminated using parity with Chunk-based similarity checking. Duplicate checking is conducted based on the file name, attribute, and size. It is needed to encrypt the files utilizing a symmetric encryption model before storing them on the cloud. Also, it removes the unwanted space and redundant data. Initially, the collected data is split into chunks and then created tags based on the chunk data. Then, it checks the weighted multi-similarity; if duplication occurs, it sends the alert; otherwise, the data is stored in the cloud. While storing the data, the optimal key-based Adaptive Elliptic Curve Cryptography (AECC) encryption is performed over the checked data for security. Here, the optimal key is created using the developed Hybrid Position of Tomtit Flock and Chameleon Swarm (HPTFCS). The result from the developed chunk-based deduplication model is evaluated through various previously used deduplication algorithms in various performance measures. The deduplication process is performed in the client side. Throughout the analysis, the statistical analysis shows 8.2% of TSA, 7.5% of GSO, 12.1% of TFOA, and 5.7% of CSO in terms of best measure. The result from the developed model achieves better performance in the deduplication process.