The task of image captioning plays a crucial role in enabling machines to describe visual content accurately and coherently. Deep models have the potential to facilitate applications as a bridge between visual content and natural language. A novel hybrid deep learning model for enhancing automated image captioning using ResNet-LSTM approach has been designed in this research. The proposed hybrid model combines the power of residual networks and long short-term memory networks to generate descriptive captions for images. To evaluate the performance of proposed model, extensive experiments were conducted and compared it to several single-stage image captioning models. Notable models included VGG16, ResNet50, and M-RNN. The results, as measured by BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores, demonstrated that the proposed hybrid model consistently outperformed the other models. Specifically, the proposed model achieved higher scores across all BLEU metrics, indicating hybrid model’s proficiency in generating contextually relevant captions. By integrating the feature extraction capabilities of ResNet with the sequential learning abilities of LSTM, the quality of generated captions was enhanced. These findings underscore the potential of hybrid models in bridging the gap between visual content and natural language. Deep learning is a transformative technology for developing applications to content understanding in diverse domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Image Captioning Through a Hybrid Deep Model: The ResNet-LSTM Approach

  • M. Lavanya,
  • B. Rupa Devi,
  • Sreenivasulu Gogula,
  • Kuntrapakam Divya,
  • Madhavi Gudavalli,
  • Shaik Mohammad Rafee

摘要

The task of image captioning plays a crucial role in enabling machines to describe visual content accurately and coherently. Deep models have the potential to facilitate applications as a bridge between visual content and natural language. A novel hybrid deep learning model for enhancing automated image captioning using ResNet-LSTM approach has been designed in this research. The proposed hybrid model combines the power of residual networks and long short-term memory networks to generate descriptive captions for images. To evaluate the performance of proposed model, extensive experiments were conducted and compared it to several single-stage image captioning models. Notable models included VGG16, ResNet50, and M-RNN. The results, as measured by BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores, demonstrated that the proposed hybrid model consistently outperformed the other models. Specifically, the proposed model achieved higher scores across all BLEU metrics, indicating hybrid model’s proficiency in generating contextually relevant captions. By integrating the feature extraction capabilities of ResNet with the sequential learning abilities of LSTM, the quality of generated captions was enhanced. These findings underscore the potential of hybrid models in bridging the gap between visual content and natural language. Deep learning is a transformative technology for developing applications to content understanding in diverse domains.