<p>The rapid advancements in the development of LLMs across diverse domains have resulted in a&#xa0;plethora of models tailored for different user groups and applications. Manual evaluation of these models poses significant challenges, including high time consumption and subjectivity. While there are established datasets for the automatic evaluation of models in English, there is a&#xa0;scarcity of datasets for other languages, such as German. A&#xa0;common approach is to translate English evaluation datasets into other languages, yet the impact of automatic translations on the functionality and comprehensibility of these datasets remains underexplored. This paper investigates the challenges of adapting datasets for the evaluation of LLMs in other languages using machine translation techniques. By employing both professional and automated translations of the MTBench dataset, we assess the effectiveness of various translation strategies using metrics such as BLEU, CHRF, and TER. Additionally, we conduct a&#xa0;qualitative analysis to identify impairments in functionality and comprehensibility caused by translations. Finally, we present examples of the results of a&#xa0;language model using the translated data sets and substantiate these with descriptive statistics. Our findings reveal significant challenges in preserving the integrity and functionality of evaluation datasets post-automatic translation, underscoring the importance of hybrid approaches that integrate both machine and human evaluations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Monolinguale Adaption von Bewertungsdatensätzen für Large Language Models: Herausforderungen und Lösungen von Übersetzungsansätzen

  • Darius Hennekeuser,
  • Daryoush Vaziri,
  • David Golchinfar,
  • Gunnar Stevens

摘要

The rapid advancements in the development of LLMs across diverse domains have resulted in a plethora of models tailored for different user groups and applications. Manual evaluation of these models poses significant challenges, including high time consumption and subjectivity. While there are established datasets for the automatic evaluation of models in English, there is a scarcity of datasets for other languages, such as German. A common approach is to translate English evaluation datasets into other languages, yet the impact of automatic translations on the functionality and comprehensibility of these datasets remains underexplored. This paper investigates the challenges of adapting datasets for the evaluation of LLMs in other languages using machine translation techniques. By employing both professional and automated translations of the MTBench dataset, we assess the effectiveness of various translation strategies using metrics such as BLEU, CHRF, and TER. Additionally, we conduct a qualitative analysis to identify impairments in functionality and comprehensibility caused by translations. Finally, we present examples of the results of a language model using the translated data sets and substantiate these with descriptive statistics. Our findings reveal significant challenges in preserving the integrity and functionality of evaluation datasets post-automatic translation, underscoring the importance of hybrid approaches that integrate both machine and human evaluations.