Vision Mamba has been growing in popularity across a variety of computer vision applications, including object detection and segmentation in remote sensing images. In order to increase the accuracy of detection and segmentation in remote sensing images, we present a unique model in this study that integrates Transformer and Vision Mamba. We employ the Mamba branch to guarantee high-quality local information extraction because of the vast scale of images and the small size of the objects. Simultaneously, we employ the Transformer branch as a residual component to the Mamba branch, with the aim of enhancing the extraction of global information. To mitigate the complexity of the Transformer branch, we employ axial attention within the self-attention block, which computes self-attention separately along vertical and horizontal directions. By integrating the Transformer and Vision Mamba, we propose the SeaMamba block and construct a novel backbone, testing it on detection and segmentation tasks on remote sensing images as well as image classification using ImageNet-1K. Our experiments demonstrate that this model achieves \(83.1\%\) accuracy in classification on the ImageNet-1K dataset. For remote sensing detection, it attains 78.12 mAP on the DOTA dataset with Oriented R-CNN detection head, and 52.32 mIOU on LoveDA dataset with MaskFormer.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-Fused Vision Mamba Model for Remote Sensing Image Detection and Segmentation

  • Zhao Yusen,
  • Wu Shan,
  • Tian Liang

摘要

Vision Mamba has been growing in popularity across a variety of computer vision applications, including object detection and segmentation in remote sensing images. In order to increase the accuracy of detection and segmentation in remote sensing images, we present a unique model in this study that integrates Transformer and Vision Mamba. We employ the Mamba branch to guarantee high-quality local information extraction because of the vast scale of images and the small size of the objects. Simultaneously, we employ the Transformer branch as a residual component to the Mamba branch, with the aim of enhancing the extraction of global information. To mitigate the complexity of the Transformer branch, we employ axial attention within the self-attention block, which computes self-attention separately along vertical and horizontal directions. By integrating the Transformer and Vision Mamba, we propose the SeaMamba block and construct a novel backbone, testing it on detection and segmentation tasks on remote sensing images as well as image classification using ImageNet-1K. Our experiments demonstrate that this model achieves \(83.1\%\) accuracy in classification on the ImageNet-1K dataset. For remote sensing detection, it attains 78.12 mAP on the DOTA dataset with Oriented R-CNN detection head, and 52.32 mIOU on LoveDA dataset with MaskFormer.