In the era of big data, enterprises are faced with the problem of ineffective data storage and management. Due to the huge amount of data and strong heterogeneity, traditional storage mechanisms often cannot meet the needs of fast access and efficient management. This paper studies an efficient big data management strategy based on distributed storage to reduce data processing latency and improve data utilization. First, by studying the characteristics of the data, a distributed storage architecture suitable for applications such as Hadoop and Spark are selected for data distribution and processing. Secondly, a data dispersion redundancy strategy and data sharding are designed to maximize data availability and security. Then, a data access mode is developed to optimize data query and data access processes while reducing data access latency. The maximum contribution of this strategy to speed can reach about 100%, and the average latency of all data types is reduced by about 8.6 ms. This big data management strategy based on distributed storage will effectively solve the pain points of enterprises in data storage and management, and lay a solid foundation for efficient data management in a big data environment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Management Strategies for Big Data Based on Distributed Storage

  • Jin Li,
  • Qingwei Guo,
  • Deming Huan

摘要

In the era of big data, enterprises are faced with the problem of ineffective data storage and management. Due to the huge amount of data and strong heterogeneity, traditional storage mechanisms often cannot meet the needs of fast access and efficient management. This paper studies an efficient big data management strategy based on distributed storage to reduce data processing latency and improve data utilization. First, by studying the characteristics of the data, a distributed storage architecture suitable for applications such as Hadoop and Spark are selected for data distribution and processing. Secondly, a data dispersion redundancy strategy and data sharding are designed to maximize data availability and security. Then, a data access mode is developed to optimize data query and data access processes while reducing data access latency. The maximum contribution of this strategy to speed can reach about 100%, and the average latency of all data types is reduced by about 8.6 ms. This big data management strategy based on distributed storage will effectively solve the pain points of enterprises in data storage and management, and lay a solid foundation for efficient data management in a big data environment.