In 2021, generative artificial intelligence technology, represented by large models, swept the world, bringing revolutionary changes to human production and life. The development of artificial intelligence has shifted from “model-centric” to “data-centric”. The theory posits that good artificial intelligence requires high-quality, large-scale, and diverse data. However, in practice, data scientists often encounter issues such as data security and privacy breaches, biased and discriminatory content output, and the problem of “high volume but low quality” data. If these issues are left unregulated, they will hinder the further development of artificial intelligence technology and may even endanger the safety of individuals, enterprises, and even national security. To address these challenges and develop more responsible and controllable artificial intelligence applications, this paper clarifies the conceptual definition of data governance for artificial intelligence (DG4AI) and proposes the main stages of artificial intelligence data governance, which are divided into nine stages including data collection and preprocessing, and the governance objects are divided into seven categories such as multimodal data and annotated data. This paper also proposes to integrate the DataOps philosophy into the data governance steps, providing enterprises with a more efficient way of data management.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Governance for Artificial Intelligence

  • Yanmei Guo,
  • Qianqian Gao,
  • Zheng Yin

摘要

In 2021, generative artificial intelligence technology, represented by large models, swept the world, bringing revolutionary changes to human production and life. The development of artificial intelligence has shifted from “model-centric” to “data-centric”. The theory posits that good artificial intelligence requires high-quality, large-scale, and diverse data. However, in practice, data scientists often encounter issues such as data security and privacy breaches, biased and discriminatory content output, and the problem of “high volume but low quality” data. If these issues are left unregulated, they will hinder the further development of artificial intelligence technology and may even endanger the safety of individuals, enterprises, and even national security. To address these challenges and develop more responsible and controllable artificial intelligence applications, this paper clarifies the conceptual definition of data governance for artificial intelligence (DG4AI) and proposes the main stages of artificial intelligence data governance, which are divided into nine stages including data collection and preprocessing, and the governance objects are divided into seven categories such as multimodal data and annotated data. This paper also proposes to integrate the DataOps philosophy into the data governance steps, providing enterprises with a more efficient way of data management.