SCAP: enhancing image captioning through lightweight feature sifting and hierarchical decoding
摘要
Image captioning aims to generate descriptive captions for visual content, thereby strengthening the connection between images and their semantic meanings. In this paper, we propose SCAP, a novel lightweight model that enhances image captioning through an innovative sifting attention mechanism. SCAP incorporates a summary module and a forget module within its encoder to refine visual information, selectively filtering visual information to retain relevant features and reduce redundancy. The hierarchical decoder then leverages sifting attention to align image features with text captions, generating accurate and contextually relevant descriptions. Extensive experiments conducted on multiple benchmark datasets, including COCO and Flickr30k, demonstrate SCAP’s effectiveness as a highly competitive model in the field. It achieves competitive performance while maintaining computational efficiency, making it particularly suitable for resource-constrained scenarios. This lightweight model represents a notable advancement in advancing image captioning.