Ara-DAQUAR: Into Arabic Question Answering on Real World Images
摘要
This paper introduces the first Arabic version of DAQUAR (DAtaset for QUestion Answering on Real-world images)) to fulfil the needs of building a visual question answering for Arabic language. A visual question answering system aims to answer an image-based question correctly. This is a multimodal task where the system process both the image and the question’s answers. Through this paper, we present the workflow of our project in which we automatically translate the English version of the DAQUAR dataset into Arabic using well-known pretrained transformers and neural machine translation toolkits. Further, we used the LaBSE model for filtering and quality assurance of the translated dataset.