Data Warehouse have become a fundamental part of organizations that want or need to leverage data. Currently, there are three main approaches for a Data Warehouse, the Inmon’s approach with a normalized model, Kimball’s with a dimensional model and Data Vault. Data Vault is a modeling technique designed to be flexible regards to changes both in the sources and analytic requirements as well as storing all versions of data for all points in time for auditing and time-travel features. This paper reviewed the latest developments in this field, those were classified in five categories, theoretical evaluation, practical evaluation, automation, metadata model and alternative models. For theoretical and practical evaluations, the papers agree on the advantages Data Vault has when it comes to model evolution and temporal aspects as well as the disadvantages on model complexity and query performance. On Data Vault automation and metadata models, the papers leverage Data Vault patterns to automate ingestion and consumption of data and model management, however it lacks standardization on those approaches. As for alternative models, those were proposed with the purpose of improve Data Vault flexibility and temporal aspects, however the resulting model is more complex in terms of number of tables and joins and the affects of this on performance, model management and automation is not covered in depth. Finally, a study of a data model that improves those aspects while reducing model complexity was proposed in which a Document Oriented and a Wide Column inspired model were evaluated against the Data Vault model, the initial results of the Wide Column inspired data model showed improved ingestion and query performance. Therefore, the next steps of the study is to further evaluate this model using larger datasets, different platforms and multiple types of analytical queries.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Data Vault Flexibility and Temporal Aspects Within the Data Warehouse Landscape

  • Victor Ruiz,
  • Luís M. Gomes,
  • Sérgio Moro

摘要

Data Warehouse have become a fundamental part of organizations that want or need to leverage data. Currently, there are three main approaches for a Data Warehouse, the Inmon’s approach with a normalized model, Kimball’s with a dimensional model and Data Vault. Data Vault is a modeling technique designed to be flexible regards to changes both in the sources and analytic requirements as well as storing all versions of data for all points in time for auditing and time-travel features. This paper reviewed the latest developments in this field, those were classified in five categories, theoretical evaluation, practical evaluation, automation, metadata model and alternative models. For theoretical and practical evaluations, the papers agree on the advantages Data Vault has when it comes to model evolution and temporal aspects as well as the disadvantages on model complexity and query performance. On Data Vault automation and metadata models, the papers leverage Data Vault patterns to automate ingestion and consumption of data and model management, however it lacks standardization on those approaches. As for alternative models, those were proposed with the purpose of improve Data Vault flexibility and temporal aspects, however the resulting model is more complex in terms of number of tables and joins and the affects of this on performance, model management and automation is not covered in depth. Finally, a study of a data model that improves those aspects while reducing model complexity was proposed in which a Document Oriented and a Wide Column inspired model were evaluated against the Data Vault model, the initial results of the Wide Column inspired data model showed improved ingestion and query performance. Therefore, the next steps of the study is to further evaluate this model using larger datasets, different platforms and multiple types of analytical queries.