AIDataSet: Knowledge Graph Construction for AI Applications
摘要
Currently, there is a lack of widely used Chinese open-source relation extraction datasets in the field of artificial intelligence (AI). To address this gap, we introduce AIDataSet, a dataset of entity relationships extracted from AI-related texts, including AI institutions, websites, companies, and research institutions. This paper details the methodology for constructing a Chinese dataset for the AI domain, which includes entity and relation annotations, aiming to provide high-quality data resources for knowledge extraction and the construction and training of domain-specific dataset models. The dataset construction process involves multiple key steps, including data acquisition, preprocessing, annotation, and augmentation. To enhance the diversity and utility of the dataset, we develop a framework called DATAEXT, which can efficiently extract triples containing entities and relations from texts. Additionally, we introduce various data augmentation methods, including distant supervision (DS) and data enhancement techniques, to generate high-quality synthetic data. Through these comprehensive methods, we construct the first Chinese dataset for the AI domain, which shows significant improvements in data diversity, accuracy, and utility across multiple dimensions. This paper expects that this systematic research will provide more high-quality and reliable data resources for the development of Chinese AI applications, thereby promoting the innovation and application expansion of related technologies.