Image Caption Generation Evaluation Based on XceptionNet, VGGNet, and ResNet
摘要
A caption generation for the image is called image captioning. Image captioning means giving the scenario of an image based on attributes presented in the image. Artificial Intelligence (AI) based technologies have been increasing rapidly for over a decade. Deep neural networks are one of the critical ingredients in AI. By using these neural networks, we are generating captions for the image. Convolutional Neural Network (CNN) is the first neural network that is used for the image caption generation and feature extraction of the image. Later, when AI- based research was increased, different types of neural networks replaced CNN and got better results. This paper aims to give the introduction of image captioning and compare the three best neural networks (XceptionNet, VGGNet, ResNet) for image captioning for an image with the BLEU scores.