Spatiotemporal Memory Network for UAV Target Tracking
摘要
Visual object tracking empowers a wide array of near-earth remote sensing applications based on unmanned aerial vehicles (UAVs). However, the frequent alterations in flight maneuvers and perspectives during UAV tracking present considerable challenges, such as target scale variation and occlusion. Existing template matching approaches often involve complex template updating mechanisms. To mitigate this limitation, we introduce a novel ViTMem network for UAV object tracking. ViTMem comprises three main modules dedicated to precise target localization. Initially, the search image and historical template images features are extracted using Vision Transformer (ViT) encoder. Subsequently, essential template features for target detection in the current frame are dynamically retrieved. Finally, composite features are processed to predict the target localization. Experiments conducted on three demanding datasets demonstrate that ViTMem surpasses state-of-the-art methods, particularly in scenarios with target scale variation and occlusion.