Improving Image Geolocation with Multimodal Deep Learning
摘要
Image geolocalization is a challenging task in computer vision that involves identifying an unknown geographic location depicted in an image. This paper presents a novel multimodal approach to the image geolocalization problem, leveraging both image and text data for accurate predictions. The proposed model utilizes high-quality street view images and textual clues from community websites to train an attention mechanism that aligns images with relevant textual information. Such an approach significantly reduces training time and the number of trainable parameters compared to existing state-of-the-art models while improving performance. The model outperforms the current best model on both a street view image test set and a benchmark dataset of social media images. The effectiveness of the proposed method is demonstrated through comprehensive analysis and explanatory examination of the results, highlighting its application potential to various real-world tasks.