<p>Hadoop Distributed File System (HDFS) is known for its specialized strategies and policies tailored to enhance replica placement. This capability is critical for ensuring efficient and reliable access to data replicas, particularly as HDFS operates best when data are evenly distributed within the cluster. In this paper, we build upon earlier practical evaluations and conduct a thorough analysis of the replica balancing process in HDFS, focusing on two critical performance metrics: stability and efficiency. We evaluated these aspects alongside balancing operational cost by contrasting them with conventional HDFS solutions and employing a novel dynamic architecture for data replica balancing. On top of that, we delve into the optimizations in data locality brought about by effective replica balancing and their benefits for data-intensive applications, including enhanced read performance. Our findings reveal the extent to which data imbalance in HDFS directly affects the file system and highlight the struggles of the default replica placement policy in maintaining cluster balance. We examined the real but intricate and temporary effectiveness of on-demand balancing, underscoring the importance of regular and adaptable balancing interventions. This reaffirms the significance of context-aware replica balancing, as provided by the proposed dynamic architecture, not only for maintaining data equilibrium but also for ensuring efficient system performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing the stability, efficiency, and cost of a dynamic data replica balancing architecture for HDFS

  • Rhauani Weber Aita Fazul,
  • Odorico Machado Mendizabal,
  • Patrícia Pitthan Barcelos

摘要

Hadoop Distributed File System (HDFS) is known for its specialized strategies and policies tailored to enhance replica placement. This capability is critical for ensuring efficient and reliable access to data replicas, particularly as HDFS operates best when data are evenly distributed within the cluster. In this paper, we build upon earlier practical evaluations and conduct a thorough analysis of the replica balancing process in HDFS, focusing on two critical performance metrics: stability and efficiency. We evaluated these aspects alongside balancing operational cost by contrasting them with conventional HDFS solutions and employing a novel dynamic architecture for data replica balancing. On top of that, we delve into the optimizations in data locality brought about by effective replica balancing and their benefits for data-intensive applications, including enhanced read performance. Our findings reveal the extent to which data imbalance in HDFS directly affects the file system and highlight the struggles of the default replica placement policy in maintaining cluster balance. We examined the real but intricate and temporary effectiveness of on-demand balancing, underscoring the importance of regular and adaptable balancing interventions. This reaffirms the significance of context-aware replica balancing, as provided by the proposed dynamic architecture, not only for maintaining data equilibrium but also for ensuring efficient system performance.