Person with impaired vision have faced issues while understanding and interacting with pictures and videos. Image captioning enhance their capacity to deal with visual content by creating descriptive language, that makes the images more accessible and engaging. This technique generates relevant captions from visual data by connecting visual feature extraction and language modeling. It acts as a key connector between computer vision and natural language processing, that allows machines to analyze and describe pictures. This article examines the developments of models, such as encoder-decoder frameworks and attention processes, that combine visual feature extraction with language creation and focusing on how object identification improves caption accuracy and contextual relevance. The study discusses on various deep learning models and also highlights various evaluation metrics like BLEU, ROUGE, and CIDEr, which measure the quality of generated captions. Moreover, the paper addresses potential improvements and prospects in image captioning systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Captioning Using Deep Learning Models: A Comprehensive Overview

  • Nisha,
  • Shailesh D. Kamble,
  • Himanshu Mittal

摘要

Person with impaired vision have faced issues while understanding and interacting with pictures and videos. Image captioning enhance their capacity to deal with visual content by creating descriptive language, that makes the images more accessible and engaging. This technique generates relevant captions from visual data by connecting visual feature extraction and language modeling. It acts as a key connector between computer vision and natural language processing, that allows machines to analyze and describe pictures. This article examines the developments of models, such as encoder-decoder frameworks and attention processes, that combine visual feature extraction with language creation and focusing on how object identification improves caption accuracy and contextual relevance. The study discusses on various deep learning models and also highlights various evaluation metrics like BLEU, ROUGE, and CIDEr, which measure the quality of generated captions. Moreover, the paper addresses potential improvements and prospects in image captioning systems.