Automatically Generating a Dataset for Natural Language Inference Systems from a Knowledge Graph
摘要
In this study, we propose a method to automatically create Vietnamese Natural Language Inference (NLI) datasets from Knowledge Graph (KG). The approach leverages information of Knowledge Graph (KG) and employs thesaurus and antonym dictionary expansion techniques to generate premise-hypothesis sentence pairs with labels of entailment, contradiction, and neutral for natural language inference (NLI). The researchers also conducted a process of validating and improving the quality of the generated dataset. The experimental results demonstrate that this method of automatically creating Vietnamese NLI datasets from KG achieves reliable and effective performance. The generated dataset not only expands the scale of existing Vietnamese NLI datasets but also provides valuable resources for training and evaluating Vietnamese NLI models. This method holds the potential for wide-ranging applications in the development of Vietnamese NLP applications and related research.