Text simplification techniques have developed significantly in recent years from rule-based to data-driven approaches. However, the lack of high-quality data sources and the need for further exploration of the low-resource scenarios make text simplification a complex and non-trivial task. In this paper, we present experiments in text simplification for Lithuanian, focusing on simplifying administrative texts to a plain language level, which is intended for the general public. We chose mT5 and mBART as foundational models and fine-tuned them for the text simplification task. We also tested ChatGPT for this task as it became popular in variety of applications. We evaluated the outputs of these models quantitatively and qualitatively to obtain a balanced assessment. All in all, mBART was found to be the most effective model for simplifying Lithuanian text, achieving the highest BLEU, ROUGE and BERTscore scores. A qualitative evaluation of the simplified sentences produced by assessing the simplicity, meaning retention and grammaticality of sentences, simplified by our fine-tuned models, added to the results of scores of evaluation metrics, though ChatGPT showed competitive results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Text Simplification for Lithuanian

  • Justina Mandravickaitė,
  • Eglė Rimkienė,
  • Danguolė Kotryna Kapkan,
  • Danguolė Kalinauskaitė,
  • Tomas Krilavičius

摘要

Text simplification techniques have developed significantly in recent years from rule-based to data-driven approaches. However, the lack of high-quality data sources and the need for further exploration of the low-resource scenarios make text simplification a complex and non-trivial task. In this paper, we present experiments in text simplification for Lithuanian, focusing on simplifying administrative texts to a plain language level, which is intended for the general public. We chose mT5 and mBART as foundational models and fine-tuned them for the text simplification task. We also tested ChatGPT for this task as it became popular in variety of applications. We evaluated the outputs of these models quantitatively and qualitatively to obtain a balanced assessment. All in all, mBART was found to be the most effective model for simplifying Lithuanian text, achieving the highest BLEU, ROUGE and BERTscore scores. A qualitative evaluation of the simplified sentences produced by assessing the simplicity, meaning retention and grammaticality of sentences, simplified by our fine-tuned models, added to the results of scores of evaluation metrics, though ChatGPT showed competitive results.