A Novel Technique for Image Captioning Based on Hierarchical Clustering and Deep Learning
摘要
The advent of deep learning has brought about significant progress in the field of computer vision, leading to its extensive application in various domains. One of the most important tasks in computer vision is image captioning, which is essential to enable machines to understand and describe visual content in natural language. Applications ranging from improving user experiences in digital media to assistive devices for the blind or visually challenged require this capability. In this paper, we introduce a novel method for accomplishing the image-captioning task with a focus on data reduction through clustering techniques. Our approach aims to maintain or surpass the accuracy achieved by prior studies. Our unique approach uses clustering algorithms to reduce data and addresses the issues of data and computing intensity. It is expected to improve the accuracy and efficiency of picture captioning models. Additionally, we investigate the outcomes of utilizing single long short-term memory (LSTM) as opposed to stacked LSTMs for generating captions. Finally, the MS-COCO benchmark image dataset is used to analyze the performance of proposed approaches, their contributions, and relevance are highlighted, emphasizing the significance of the proposed approach and shedding light on a potential research direction for the image captioning approaches.