Evaluating the Performance of LLMs When Translating Saudi Arabic as Low Resource Language
摘要
This paper evaluates the performance of different large language models (LLMs) in translating textual data from Saudi Arabic, a low-resource language, into English. In this investigation we employ the state-of-the-art language models namely; ChatGPT-4, Claude-3 and Palm-2. We assess the capabilities of these LLMs on the Arabic Semantic Textual Similarity (STS) dataset. The evaluation covers different aspects, including the standard evaluation metrics, prompt design, and comparison with baselines systems namely; Google Translator, QuillBot Translator and Systran Translator. We conducted human evaluation on the generated translation and analysis the most frequent translation error using our sample dataset and different models. Our findings reveal significant insights into the strengths of ChatGPT (GPT-4) model in handling and translating dialectal Arabic with the highest Bilingual Evaluation Understudy (BLEU) score among all participated models (46.56).