TMCL-GEO: Text-Guided Multimodal Contrastive Learning for Cross-View Geo-Localization
摘要
Cross-view geo-localization aims to align images captured from ground cameras or drones with those from satellite views, enabling geographic localization using GPS-enriched satellite imagery. Most existing methodologies predominantly rely on extracting consistent features through cross-view contrastive learning. However, the nearly orthogonal relationship between ground and satellite views presents significant challenges in effectively learning consistent information solely from images. To address this, we propose a multimodal framework that incorporates textual information. By achieving multi-level alignment between text and image modalities, the framework leverages the invariance of textual information to facilitate the learning of cross-view invariant features. Our approach has been validated on cross-view localization datasets, demonstrating state-of-the-art performance.