<p>Benchmark corpora for translation are crucial for assessing machine translation (MT) systems. These corpora allow researchers and developers to evaluate the performance, accuracy, and efficiency of their translation models. In this context, our work introduces a new translation benchmark <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\hbox {corpus}^{*}\)</EquationSource> </InlineEquation> for translations between Indian languages, with Hindi as the source language, spanning 12 languages (Assamese, Bangla, English, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, and Urdu) in the governance domain. Translations are sourced online, automatically aligned, and then undergo human validation and correction for alignment and translation. Additional human validation, further refines these translations for benchmark use. The corpus comprises 1304 n-way parallel sentences and three bilingual sets (dev, devtest, and test), each containing 1000 to 3000 sentences, with Hindi serving as the source language. We offer these corpus to the research community for testing and evaluating MT systems in the governance domain and between Indian languages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

IL-ILGOV-2024: a translation benchmark for Hindi-to-12 languages in the governance domain

  • Vandan Mujadia,
  • Rao B. Ashwath,
  • Dipti Misra Sharma

摘要

Benchmark corpora for translation are crucial for assessing machine translation (MT) systems. These corpora allow researchers and developers to evaluate the performance, accuracy, and efficiency of their translation models. In this context, our work introduces a new translation benchmark \(\hbox {corpus}^{*}\) for translations between Indian languages, with Hindi as the source language, spanning 12 languages (Assamese, Bangla, English, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, and Urdu) in the governance domain. Translations are sourced online, automatically aligned, and then undergo human validation and correction for alignment and translation. Additional human validation, further refines these translations for benchmark use. The corpus comprises 1304 n-way parallel sentences and three bilingual sets (dev, devtest, and test), each containing 1000 to 3000 sentences, with Hindi serving as the source language. We offer these corpus to the research community for testing and evaluating MT systems in the governance domain and between Indian languages.