Evaluation of Large Language Models for the Vietnamese Language in Generative Vietnamese Economy Chatbots (GVEC) Services
摘要
The demand for automated question-answering capabilities in chatbot services has become increasingly critical, influencing various aspects of daily life. This study focuses on developing a chatbot system specialized in economic topics within Vietnam. We compare several Vietnamese-compatible large language models within a retrieval-augmented generation framework for question-answering. Our new benchmark, based on the Vietnamese Economy Information Database, includes 1000 question-answer pairs—a substantial increase from previous studies, which relied on a smaller dataset of 100 pairs. Earlier studies were limited by the high cost of ChatGPT testing and focused on exact data accuracy rather than evaluating overall response quality. These studies only compared GPT-3.5-turbo with three more affordable, open-source models to identify viable alternatives. In contrast, our study’s larger dataset and new experiments on the updated GPT-4o mini, along with additional evaluation metrics, enable a more comprehensive assessment, addressing both response accuracy and overall answer quality.