With the rapid development of large language models, a lot of preliminary work has been done to explore the performance of large language models on different artificial intelligence tasks. Despite this, research on evaluating the translation performance of large language models for Southeast Asian low-resource languages is still relatively scarce. The main reason is the scarcity of public evaluation data for Southeast Asian low-resource languages and large language models that can be directly injected into Southeast Asian low-resource languages for training. First, evaluation data for Southeast Asian low-resource language is extracted based on the publicly available Asian Language Treebank (ALT) corpus. Then, large language models with the capability to understand Southeast Asian low-resource languages are evaluated. Finally, three strategies are employed to construct in-context learning prompts to further enhance the translation performance of the existing large language models. Experimental results on multiple benchmark datasets demonstrate that large language models have superior contextual learning capabilities and their translation performances can be improved significantly by utilizing high-quality prompts. However, the guidance effect of sub-optimal prompts on the model is inconsistent. The comparative experiment further elucidated that the model temperature parameter exhibits distinct optimal values depending on the scale of text input. The subsequent analysis indicates that a higher degree of similarity between the prompt and the context facilitates the model’s ability to generate accurate translation results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Translation Performance of Multilingual Large Language Models: A Case Study on Southeast Asian Languages

  • Hua Lai,
  • Silei Han,
  • Ying Li,
  • Zhengtao Yu,
  • Zirui Guo

摘要

With the rapid development of large language models, a lot of preliminary work has been done to explore the performance of large language models on different artificial intelligence tasks. Despite this, research on evaluating the translation performance of large language models for Southeast Asian low-resource languages is still relatively scarce. The main reason is the scarcity of public evaluation data for Southeast Asian low-resource languages and large language models that can be directly injected into Southeast Asian low-resource languages for training. First, evaluation data for Southeast Asian low-resource language is extracted based on the publicly available Asian Language Treebank (ALT) corpus. Then, large language models with the capability to understand Southeast Asian low-resource languages are evaluated. Finally, three strategies are employed to construct in-context learning prompts to further enhance the translation performance of the existing large language models. Experimental results on multiple benchmark datasets demonstrate that large language models have superior contextual learning capabilities and their translation performances can be improved significantly by utilizing high-quality prompts. However, the guidance effect of sub-optimal prompts on the model is inconsistent. The comparative experiment further elucidated that the model temperature parameter exhibits distinct optimal values depending on the scale of text input. The subsequent analysis indicates that a higher degree of similarity between the prompt and the context facilitates the model’s ability to generate accurate translation results.