With the rapid development of satellite and remote sensing technologies, remote sensing images are practically valuable for many fields in practical applications. Especially in object detection tasks, how to effectively simulate and exploit long range dependencies in images is a key challenge. In this paper, a novel Full-Scale Network (FSNet) is proposed, where the projection operations are no longer required. Firstly, we introduce a Intra-Scale Feature Enhancement (ISFE) module to extract rich spatial details by selective scan, which enables intra-scale feature fusion for long-range dependencies without projection. Secondly, we develop a Cross-Scale Detection Head (CSDH) to model the long-range dependencies between different scales by jointly up-sampling and down- sampling the output of the ISFE module. In this way, the back-projection is also avoided. The experimental results on the NWPU VHR-10 and DOTA datasets verify the effectiveness of this method. The mAP of our method is 1.3% better than the best methods in the past.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Full-Scale Network for Remote Sensing Object Detection

  • Xiaofeng Pan,
  • Xiang Zhang,
  • Peng Wang,
  • Jianan Hou

摘要

With the rapid development of satellite and remote sensing technologies, remote sensing images are practically valuable for many fields in practical applications. Especially in object detection tasks, how to effectively simulate and exploit long range dependencies in images is a key challenge. In this paper, a novel Full-Scale Network (FSNet) is proposed, where the projection operations are no longer required. Firstly, we introduce a Intra-Scale Feature Enhancement (ISFE) module to extract rich spatial details by selective scan, which enables intra-scale feature fusion for long-range dependencies without projection. Secondly, we develop a Cross-Scale Detection Head (CSDH) to model the long-range dependencies between different scales by jointly up-sampling and down- sampling the output of the ISFE module. In this way, the back-projection is also avoided. The experimental results on the NWPU VHR-10 and DOTA datasets verify the effectiveness of this method. The mAP of our method is 1.3% better than the best methods in the past.