This paper introduces the ViFoodNLI dataset, a natural language inference (NLI) dataset for Vietnamese. While recent efforts have been made to build high-quality NLI datasets for Vietnamese and some Cross-Lingual NLI Corpus (with support for Vietnamese) for multiple domains, our dataset specifically focuses on the field of local cuisine. The main reason for choosing this field is that cuisine is a significant component of Vietnamese culture, and thus the dataset encompasses many characteristics of the Vietnamese language. By collecting information on culinary topics from reliable news sources, we have developed various methods and logics such as knowledge graphs and Generative AI to create high-quality pairs of premise and hypothesis sentences. Through rigorous testing, the dataset has achieved significant results, creating momentum for future research and practical applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViFoodNLI: A Dataset for Vietnamese Natural Language Inference in Local Cuisine

  • Long Ngo Hoang Phan,
  • Phuc Do

摘要

This paper introduces the ViFoodNLI dataset, a natural language inference (NLI) dataset for Vietnamese. While recent efforts have been made to build high-quality NLI datasets for Vietnamese and some Cross-Lingual NLI Corpus (with support for Vietnamese) for multiple domains, our dataset specifically focuses on the field of local cuisine. The main reason for choosing this field is that cuisine is a significant component of Vietnamese culture, and thus the dataset encompasses many characteristics of the Vietnamese language. By collecting information on culinary topics from reliable news sources, we have developed various methods and logics such as knowledge graphs and Generative AI to create high-quality pairs of premise and hypothesis sentences. Through rigorous testing, the dataset has achieved significant results, creating momentum for future research and practical applications.