<p>Image captioning generates descriptive text from an input image, establishing a connection between the image content and words. Recently, the most successful approaches for automatically creating image captions have been based on transformer learning models. Arabic image captioning has gained importance due to the unique characteristics of the Arabic language. This paper introduces an attention-based transformer model for Arabic image captioning (ARTIC). ARTIC employs a deep learning convolutional neural network (CNN) for feature extraction from the images and a transformer encoder–decoder architecture for generating textual captions. ARTIC utilizes an ensemble learning approach based on a voting mechanism that selects the caption with the highest bilingual evaluation understudy (BLEU) score to produce the captions. To evaluate the effectiveness of the proposed model, the publicly available Flickr8k benchmark dataset was used for Arabic image captioning. Our results show that ARTIC achieved the best scores for BLEU-1, CIDEr, and ROUGE at rates of (0.626), (0.838), and (0.471), respectively. The other metrics, such as BLEU-2 and METEOR, achieved competitive rates of (0.381) and (0.332), respectively. The experiments with the Flickr30k English dataset demonstrated the generalizability of the proposed approach to other languages. These results indicate that the suggested model outperformed other models used for comparison. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-based transformer model for Arabic image captioning

  • Israa Al Badarneh,
  • Rana Husni Al Mahmoud,
  • Bassam H. Hammo,
  • Omar Al-Kadi

摘要

Image captioning generates descriptive text from an input image, establishing a connection between the image content and words. Recently, the most successful approaches for automatically creating image captions have been based on transformer learning models. Arabic image captioning has gained importance due to the unique characteristics of the Arabic language. This paper introduces an attention-based transformer model for Arabic image captioning (ARTIC). ARTIC employs a deep learning convolutional neural network (CNN) for feature extraction from the images and a transformer encoder–decoder architecture for generating textual captions. ARTIC utilizes an ensemble learning approach based on a voting mechanism that selects the caption with the highest bilingual evaluation understudy (BLEU) score to produce the captions. To evaluate the effectiveness of the proposed model, the publicly available Flickr8k benchmark dataset was used for Arabic image captioning. Our results show that ARTIC achieved the best scores for BLEU-1, CIDEr, and ROUGE at rates of (0.626), (0.838), and (0.471), respectively. The other metrics, such as BLEU-2 and METEOR, achieved competitive rates of (0.381) and (0.332), respectively. The experiments with the Flickr30k English dataset demonstrated the generalizability of the proposed approach to other languages. These results indicate that the suggested model outperformed other models used for comparison.