<p>This paper introduces a solution to improve the quality of Vietnamese text summarization by utilizing the Vistral 7B large language model (LLM) that specializes in supporting Vietnamese, specifically through fine-tuning combined with the QDoRA (Quantized Decomposed Low-Rank Adaptation) technique. Central to our methodology is a rigorous data filtering strategy applied to the training corpus, designed to refine the dataset and ensure that the resulting summaries exhibit high fidelity. At the same time, using QDoRA solves the important problem of memory limits that often come with large models. This advanced quantization method dramatically reduces the model’s memory footprint without causing a significant drop in performance. In addition, we also implement advanced optimization techniques from DeepSpeed, enabling efficient and scalable training of large language models (LLMs). The Vistral 7B model is a training version of the original Mistral 7B model. It has been specifically trained on curated Vietnamese datasets to ensure relevance and diversity of the data. We assessed the efficacy of Vistral 7B fine-tuned with QDoRA by comparing its performance against the ViT5 model, which was designed for Vietnamese text summarization. Experimental data demonstrate that our proposed method is better than ViT5, and this is most clearly seen in the results of the ROUGE evaluations. Specifically, on the preprocessed dataset, our method achieved a ROUGE-1 score of 76.42, a ROUGE-2 score of 57.88, and a ROUGE-L score of 59.00, compared to ViT5<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(_{base} \)</EquationSource> </InlineEquation>’s scores of 71.85, 51.95, and 53.40, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Resource-Efficient Vietnamese Text Summarization: Enhancing Vistral 7B Performance Through Data Filtering, QDoRA’s Low-Memory Footprint, and DeepSpeed’s Training Optimization

  • Huy Duc Nguyen Pham,
  • Dang Tuan Nguyen

摘要

This paper introduces a solution to improve the quality of Vietnamese text summarization by utilizing the Vistral 7B large language model (LLM) that specializes in supporting Vietnamese, specifically through fine-tuning combined with the QDoRA (Quantized Decomposed Low-Rank Adaptation) technique. Central to our methodology is a rigorous data filtering strategy applied to the training corpus, designed to refine the dataset and ensure that the resulting summaries exhibit high fidelity. At the same time, using QDoRA solves the important problem of memory limits that often come with large models. This advanced quantization method dramatically reduces the model’s memory footprint without causing a significant drop in performance. In addition, we also implement advanced optimization techniques from DeepSpeed, enabling efficient and scalable training of large language models (LLMs). The Vistral 7B model is a training version of the original Mistral 7B model. It has been specifically trained on curated Vietnamese datasets to ensure relevance and diversity of the data. We assessed the efficacy of Vistral 7B fine-tuned with QDoRA by comparing its performance against the ViT5 model, which was designed for Vietnamese text summarization. Experimental data demonstrate that our proposed method is better than ViT5, and this is most clearly seen in the results of the ROUGE evaluations. Specifically, on the preprocessed dataset, our method achieved a ROUGE-1 score of 76.42, a ROUGE-2 score of 57.88, and a ROUGE-L score of 59.00, compared to ViT5 \(_{base} \) ’s scores of 71.85, 51.95, and 53.40, respectively.