Hajj-FQA: A benchmark Arabic dataset for developing question-answering systems on Hajj fatwas
摘要
Deep learning has significantly advanced the question-answering (QA) systems across various sectors. However, Arabic-language systems for Hajj-related fatwas (non-binding Islamic legal opinions issued by muftis) remain underdeveloped. This paper introduces Hajj-FQA, a benchmark Arabic dataset specifically designed to develop HajjBot - a specialized chatbot for fatwas QA during the annual Hajj pilgrimage. The dataset captures the unique linguistic and jurisprudential characteristics of pilgrims’ inquiries, enabling accurate, domain-specific responses. We present a comprehensive quantitative analysis of the dataset’s construction methodology and its distinctive question-answer patterns. Evaluation using multilingual and Arabic-specific language models across three tasks - machine reading comprehension (MRC), duplicate question detection (DQD), and duplicate answer detection (DAD) - with 10-fold cross-validation demonstrates the practical utility of Hajj-FQA. Results show exceptional performance in classification tasks (AraBERTv0.2 achieved a precision score of 99.19% for DQD and 99.26% for DAD) and strong extractive answering capability with an