Passage ranking is arguably one of the most important natural text-processing tasks. It is based on retrieving a passage from a larger corpus that contains an answer to a specific question. This task has not been fully explored in the Classical Arabic (CA) language. Because the size of the only available CA corpus was small, we created a new, larger CA dataset. This dataset was used to train eight Arabic pre-trained transformer-based models, which had been used in different approaches to passage ranking. The highest result we achieved was in using AraBERT Base (Arabic bidirectional encoder representations from transformers) as a bi-encoder model in the dense passage retrieval approach. This model outperformed BM25 (best match 25), which is considered a traditional approach. It obtained a 0.244 mean average precision (MAP) and 0.413 mean reciprocal rank (MRR), compared to the 0.170 MAP and 0.313 MRR obtained by BM25.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Qur’an Passage Ranking Using Transformer Models

  • Sarah Alnefaie,
  • Eric Atwell,
  • Mohammed Ammar Alsalka

摘要

Passage ranking is arguably one of the most important natural text-processing tasks. It is based on retrieving a passage from a larger corpus that contains an answer to a specific question. This task has not been fully explored in the Classical Arabic (CA) language. Because the size of the only available CA corpus was small, we created a new, larger CA dataset. This dataset was used to train eight Arabic pre-trained transformer-based models, which had been used in different approaches to passage ranking. The highest result we achieved was in using AraBERT Base (Arabic bidirectional encoder representations from transformers) as a bi-encoder model in the dense passage retrieval approach. This model outperformed BM25 (best match 25), which is considered a traditional approach. It obtained a 0.244 mean average precision (MAP) and 0.413 mean reciprocal rank (MRR), compared to the 0.170 MAP and 0.313 MRR obtained by BM25.