Recent studies dealing with Abstractive Summarization are dominated by the use of Pre-trained Language Models based on Transformers. While the main contributions are applied to English, a review of the literature highlights the existence of a trend towards applying this framework on Arabic. This paper describes the full pipeline of Fine-tuning a Pre-trained Language Model based on Transformers for Arabic Abstractive Summarization. The model used is AraBART. The experiments are conducted on AHS dataset. Our work also challenges the quality of this dataset regarding the effects of repetitive summaries on the performances of the model. We found that their effect is substantial pointing out the need of a thorough study to be conducted on this dataset. A score of 54.69 \(ROUGE_1\) is obtained on the test dataset. This score drops to 46.32 when the repetitive summaries are removed. A detailed analysis is provided discussing this issue.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Tuning AraBART on AHS Dataset for Arabic Abstractive Summarization

  • Mustapha Benbarka,
  • Moulay Abdellah Kassimi

摘要

Recent studies dealing with Abstractive Summarization are dominated by the use of Pre-trained Language Models based on Transformers. While the main contributions are applied to English, a review of the literature highlights the existence of a trend towards applying this framework on Arabic. This paper describes the full pipeline of Fine-tuning a Pre-trained Language Model based on Transformers for Arabic Abstractive Summarization. The model used is AraBART. The experiments are conducted on AHS dataset. Our work also challenges the quality of this dataset regarding the effects of repetitive summaries on the performances of the model. We found that their effect is substantial pointing out the need of a thorough study to be conducted on this dataset. A score of 54.69 \(ROUGE_1\) is obtained on the test dataset. This score drops to 46.32 when the repetitive summaries are removed. A detailed analysis is provided discussing this issue.