Monolinguale Adaption von Bewertungsdatensätzen für Large Language Models: Herausforderungen und Lösungen von Übersetzungsansätzen
摘要
The rapid advancements in the development of LLMs across diverse domains have resulted in a plethora of models tailored for different user groups and applications. Manual evaluation of these models poses significant challenges, including high time consumption and subjectivity. While there are established datasets for the automatic evaluation of models in English, there is a scarcity of datasets for other languages, such as German. A common approach is to translate English evaluation datasets into other languages, yet the impact of automatic translations on the functionality and comprehensibility of these datasets remains underexplored. This paper investigates the challenges of adapting datasets for the evaluation of LLMs in other languages using machine translation techniques. By employing both professional and automated translations of the MTBench dataset, we assess the effectiveness of various translation strategies using metrics such as BLEU, CHRF, and TER. Additionally, we conduct a qualitative analysis to identify impairments in functionality and comprehensibility caused by translations. Finally, we present examples of the results of a language model using the translated data sets and substantiate these with descriptive statistics. Our findings reveal significant challenges in preserving the integrity and functionality of evaluation datasets post-automatic translation, underscoring the importance of hybrid approaches that integrate both machine and human evaluations.