Image Captioning with Neural Style Transfer Using GPT-2 and Vision Transformer Architectures
摘要
The technique NST is used to convert one image into another image without changing of the content. The only change is the configurations of the image. The content describes the layout and style representing the paint or the colors. NST deals with Content and Style images. This research paper presents a novel integration of neural style transfer and image captioning, combining artistic expression with semantic understanding. Vision Transformers and GPT-2 are two transformer-based language models that have shown extraordinary proficiency in comprehending and producing genuine language. The “Fast Arbitrary Image Styling” model from TensorFlow Hub gives users to combine multiple styles and shows various artistic elements. To obtain the model’s capability to produce captions for stylized images fosters enhanced user interaction, facilitating deeper connections between artists and the model’s outputs. This fusion of artistic flair and semantic understanding offers promising potential for innovative and expressive artistic endeavours.