Occlusion-aware dual-memory transformer for unsupervised UAV video stabilization with adaptive inference paths
摘要
Stabilizing unmanned aerial vehicle (UAV) videos remains challenging due to severe jitter, occlusion, and nonlinear scene motion. Classical trajectory-smoothing methods often assume rigid motion and fail under parallax and occlusion, while recent learning-based stabilizers such as DeepStab and DUT apply fixed pipelines that ignore frame-level variations in jitter and visibility. Transformer-based frameworks have improved temporal modeling, yet they still lack explicit occlusion reasoning and adaptive computation, limiting robustness and efficiency in dynamic UAV environments. We propose the Occlusion-Aware Dual-Memory Transformer (OADMT), an unsupervised framework that separates motion and occlusion cues via dual memory branches. An occlusion-gated attention mechanism suppresses unreliable regions, while a multi-resolution decoder and uncertainty-guided inpainting enhance stabilization quality. An adaptive inference path selection module dynamically switches between lightweight and full-capacity branches based on jitter and occlusion scores, balancing accuracy with computational cost. Experiments on UAV123-OS, DeepStab, and synthetic UAV datasets demonstrate that OADMT reduces inter-frame jitter by 17.6%, improves PSNR by 12.3%, and lowers inference time by 21.4% compared with DeepStab and other recent stabilizers, while maintaining SSIM above 93% in cross-domain tests. These results establish OADMT as an efficient and occlusion-robust stabilizer well suited for real-time UAV applications.