External Validation of Paraphrasing Language Models with Respect to Bulgarian
摘要
This paper presents the results of an external validation study evaluating five large language models - GPT-3, GPT-4, PaLM 2, Gemini 1.0 Pro, and BgGPT - on their ability to paraphrase Bulgarian sentences. The methodology developed for this research is thoroughly outlined. The study justifies the selection of two multilingual datasets containing Bulgarian, TaPaCo and XNLI, and explains the choice of the five major language models, along with the four widely-used similarity metrics for text generation evaluation: BLEU, METEOR, ROUGE, and WER. Additionally, the paper provides an analysis of the experimental results using those metrics and comparison with human evaluation. It draws conclusions related to the performance of the selected large language models in paraphrasing in Bulgarian.