In this paper, we propose a new data model capable of helping to solve the problem of data mining and data quality, particularly in Big Data. The latter is a real challenge due to the structural and semantic variations in optimal warehouse management since the development of modern digital processing techniques. Source certainty plays an important role in the semi-structured data warehouse. In addition, for correspondence, the ETL (Extraction, Transformation and Loading) process must be mastered to reduce heterogeneity. The question is how to deploy Ant Colony (ACO) and Gradient Boosting Machine (GBM) algorithms. Then combine these for semi-structured reference data. And finally, make them efficient for schema enhancement. Experiments show that the hybrid approach outperforms the individual algorithms in terms of precision (98.9%), recall (97.8%) and F1-measurement (97.2%). This method is therefore significantly efficient in terms of scheme flexibility, while minimizing storage and processing costs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A New Approach to Matching Big Data Sources in a Semi-structured Data Warehouse Context

  • Anassin Chiatsè Mireille Patricia,
  • Tounwendyam Frédéric Ouédraogo,
  • Saha Kouassi Bernard

摘要

In this paper, we propose a new data model capable of helping to solve the problem of data mining and data quality, particularly in Big Data. The latter is a real challenge due to the structural and semantic variations in optimal warehouse management since the development of modern digital processing techniques. Source certainty plays an important role in the semi-structured data warehouse. In addition, for correspondence, the ETL (Extraction, Transformation and Loading) process must be mastered to reduce heterogeneity. The question is how to deploy Ant Colony (ACO) and Gradient Boosting Machine (GBM) algorithms. Then combine these for semi-structured reference data. And finally, make them efficient for schema enhancement. Experiments show that the hybrid approach outperforms the individual algorithms in terms of precision (98.9%), recall (97.8%) and F1-measurement (97.2%). This method is therefore significantly efficient in terms of scheme flexibility, while minimizing storage and processing costs.