Image Captioning Using Attention Mechanisms
摘要
There have been intensive developments in the domain of image caption generation using deep neural networks and attention mechanisms. This paper attempts to provide the readers a comprehensive research on the various techniques possible for image caption generation models and multiple attention mechanisms that can be incorporated within these models to significantly enhance the results. The paper explains the workflow of a general image captioning model along with the different techniques that can be leveraged within each step (feature extraction, visual encoding, language generation). This paper also reviews the standard datasets that are used in image captioning models along with the metrics that can be used for the evaluation of these models.