Empowering Image Captioning on Indian Remote Sensing Imagery Using Deep Learning
摘要
This research explores the uses of image captioning techniques in the context of remote sensing, with a particular emphasis on Indian remote sensing acquired through ISRO’s IRS-R2A satellite and LISS-IV sensors. We present a meticulously curated dataset encompassing diverse landscapes across Indian states, ensuring a comprehensive representation of the distinct characteristics inherent in satellite imagery. Through a systematic evaluation process involving various CNN models in conjunction with LSTM and Vision Transformers, we determine the most effective model based on various evaluation metrics. Notably, our fusion which combines ENet-B0 and ViT, achieves a remarkable accuracy rate of 90.63%. This study makes a valuable contribution to the advancement of image captioning in remote sensing areas, offering potential applications in fields such as disaster management and urban planning.