Video object detection is a challenging task because the complex motion of drones will lead to a decline in visual effects. A typical solution is to aggregate adjacent features to enhance the appearance features of each frame. However, direct aggregation not only introduces a large amount of background, but also ignores the rich motion information in the video. Aiming at the problem of appearance degradation in UAV videos, a Spatio-Temporal Correlation (STC) network is proposed to mine time context information through cross-view information exchange to achieve video target detection. Instead of using Transformer’s attention mechanism, a spatio-temporal extraction module with linear computational complexity and cross-view perception ability is proposed to selectively aggregate features from other frames. Our STC network achieves excellent performance on the VisDrone-VID dataset and has a faster runtime. Specifically, our STC-YOLOv7 achieved 24.7% mAP (6% higher than YOLOv7) and 37 Frames Per Second (FPS).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Improved Algorithm for UAV Video Object Detection Based on Spatiotemporal Correlation

  • Zihao Zhou,
  • Xianguo Yu,
  • Xiangke Wang

摘要

Video object detection is a challenging task because the complex motion of drones will lead to a decline in visual effects. A typical solution is to aggregate adjacent features to enhance the appearance features of each frame. However, direct aggregation not only introduces a large amount of background, but also ignores the rich motion information in the video. Aiming at the problem of appearance degradation in UAV videos, a Spatio-Temporal Correlation (STC) network is proposed to mine time context information through cross-view information exchange to achieve video target detection. Instead of using Transformer’s attention mechanism, a spatio-temporal extraction module with linear computational complexity and cross-view perception ability is proposed to selectively aggregate features from other frames. Our STC network achieves excellent performance on the VisDrone-VID dataset and has a faster runtime. Specifically, our STC-YOLOv7 achieved 24.7% mAP (6% higher than YOLOv7) and 37 Frames Per Second (FPS).