<p>Stabilizing unmanned aerial vehicle (UAV) videos remains challenging due to severe jitter, occlusion, and nonlinear scene motion. Classical trajectory-smoothing methods often assume rigid motion and fail under parallax and occlusion, while recent learning-based stabilizers such as DeepStab and DUT apply fixed pipelines that ignore frame-level variations in jitter and visibility. Transformer-based frameworks have improved temporal modeling, yet they still lack explicit occlusion reasoning and adaptive computation, limiting robustness and efficiency in dynamic UAV environments. We propose the Occlusion-Aware Dual-Memory Transformer (OADMT), an unsupervised framework that separates motion and occlusion cues via dual memory branches. An occlusion-gated attention mechanism suppresses unreliable regions, while a multi-resolution decoder and uncertainty-guided inpainting enhance stabilization quality. An adaptive inference path selection module dynamically switches between lightweight and full-capacity branches based on jitter and occlusion scores, balancing accuracy with computational cost. Experiments on UAV123-OS, DeepStab, and synthetic UAV datasets demonstrate that OADMT reduces inter-frame jitter by 17.6%, improves PSNR by 12.3%, and lowers inference time by 21.4% compared with DeepStab and other recent stabilizers, while maintaining SSIM above 93% in cross-domain tests. These results establish OADMT as an efficient and occlusion-robust stabilizer well suited for real-time UAV applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Occlusion-aware dual-memory transformer for unsupervised UAV video stabilization with adaptive inference paths

  • Roshnadevi Jaising Sapkal,
  • Krishna K. Warhade

摘要

Stabilizing unmanned aerial vehicle (UAV) videos remains challenging due to severe jitter, occlusion, and nonlinear scene motion. Classical trajectory-smoothing methods often assume rigid motion and fail under parallax and occlusion, while recent learning-based stabilizers such as DeepStab and DUT apply fixed pipelines that ignore frame-level variations in jitter and visibility. Transformer-based frameworks have improved temporal modeling, yet they still lack explicit occlusion reasoning and adaptive computation, limiting robustness and efficiency in dynamic UAV environments. We propose the Occlusion-Aware Dual-Memory Transformer (OADMT), an unsupervised framework that separates motion and occlusion cues via dual memory branches. An occlusion-gated attention mechanism suppresses unreliable regions, while a multi-resolution decoder and uncertainty-guided inpainting enhance stabilization quality. An adaptive inference path selection module dynamically switches between lightweight and full-capacity branches based on jitter and occlusion scores, balancing accuracy with computational cost. Experiments on UAV123-OS, DeepStab, and synthetic UAV datasets demonstrate that OADMT reduces inter-frame jitter by 17.6%, improves PSNR by 12.3%, and lowers inference time by 21.4% compared with DeepStab and other recent stabilizers, while maintaining SSIM above 93% in cross-domain tests. These results establish OADMT as an efficient and occlusion-robust stabilizer well suited for real-time UAV applications.