A Fusion-Based Approach for Generating Image Captions
摘要
Image captioning is a demanding task that involves generating textual descriptions of images. In recent years, deep learning approaches, such as convolution neural networks (CNNs) and long short-term memory (LSTM) networks, have been extensively used to address this task. In this approach, the CNN is used to extract features from the input image, while the LSTM is used to produce a succession of words that describe the image. The CNN-LSTM architecture has been revealed to achieve modern performance on various image captioning datasets. In this chapter, we present a review of the current progress in image captioning using CNN and LSTM, including the different architectures and techniques used in the literature. Adding the capability to read aloud the image subtitle has increased the sophistication of image captioning in comparison to other works in the field.