<p>In the cloud environment, number of clients are generating huge amounts of duplicate content, in the form of different types of files. Data redundancy is more prevalent among the same type of files, and it is negligible across different types of files. Applying the same deduplication technique irrespective of the file type results in wastage of resources, with insignificant duplicates elimination. If the same type of files is routed to the same data server, with deduplication, more duplicate content can be eliminated, but raises the storage imbalance problem. If files are distributed among data servers, ignoring the file type, better storage balance can be achieved, but the eliminated duplicate content decreases. In order to address these challenges, a distributed deduplication system is proposed. The system categorizes files into three groups—high, low, and unpredictable duplicate files—based on their redundancy levels. A group of data servers is allocated for each category of files. Preprocessing at source and resource-intensive duplicate content identification and elimination at the data servers reduces communication overhead. Data servers apply similarity-based segment-level deduplication for high and unpredictable duplicate files, and file-level deduplication is applied for low duplicate files. This system gives a high degree of load balance while achieving comparable space saving.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

File aware distributed deduplication system in cloud environment

  • Amdewar Godavari,
  • Chapram Sudhakar,
  • T. Ramesh

摘要

In the cloud environment, number of clients are generating huge amounts of duplicate content, in the form of different types of files. Data redundancy is more prevalent among the same type of files, and it is negligible across different types of files. Applying the same deduplication technique irrespective of the file type results in wastage of resources, with insignificant duplicates elimination. If the same type of files is routed to the same data server, with deduplication, more duplicate content can be eliminated, but raises the storage imbalance problem. If files are distributed among data servers, ignoring the file type, better storage balance can be achieved, but the eliminated duplicate content decreases. In order to address these challenges, a distributed deduplication system is proposed. The system categorizes files into three groups—high, low, and unpredictable duplicate files—based on their redundancy levels. A group of data servers is allocated for each category of files. Preprocessing at source and resource-intensive duplicate content identification and elimination at the data servers reduces communication overhead. Data servers apply similarity-based segment-level deduplication for high and unpredictable duplicate files, and file-level deduplication is applied for low duplicate files. This system gives a high degree of load balance while achieving comparable space saving.