<p>Remote sensing images (RSIs) inherently contain rich spectral information across different frequency bands, yet current semantic segmentation networks predominantly focus on spatial feature extraction while neglecting frequency domain analysis, leading to suboptimal segmentation performance. Spectra in different frequency bands and at different scales exhibit complementarity. Reasonable frequency representation can effectively capture inter-category discriminative patterns that are imperceptible in spatial domain. This paper presents a dual-domain learning framework (FCTNet) that synergistically integrates frequency information with convolutional neural network (CNN) and Transformer architecture, enabling comprehensive feature extraction in both spatial and frequency domains. Our encoder employs a hybrid approach combining fast Fourier convolution with central difference convolution, achieving expanded receptive fields while preserving structural details. The decoder features a concurrent interaction architecture comprising a dynamic multi-scale spatial convolution branch and a dual-scale frequency-adjusted Transformer branch, which establish inter-domain communication through cross-attention mechanisms. Notably, the frequency-adjusted Transformer incorporates adaptive frequency allocation that dynamically prioritizes critical frequency components while suppressing non-essential frequency constituents. This mechanism help improve segmentation accuracy and model robustness. Experiments on two benchmark RSI segmentation datasets demonstrate the superior performance of our method, with ablation studies validating the effectiveness of each proposed component.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fourier aids CNN and transformer for semantic segmentation of remote sensing images

  • Jun Wang,
  • Youzhou Wu,
  • Baodi Liu,
  • Wenzheng Wang,
  • Haoran Xu,
  • Keding Wang

摘要

Remote sensing images (RSIs) inherently contain rich spectral information across different frequency bands, yet current semantic segmentation networks predominantly focus on spatial feature extraction while neglecting frequency domain analysis, leading to suboptimal segmentation performance. Spectra in different frequency bands and at different scales exhibit complementarity. Reasonable frequency representation can effectively capture inter-category discriminative patterns that are imperceptible in spatial domain. This paper presents a dual-domain learning framework (FCTNet) that synergistically integrates frequency information with convolutional neural network (CNN) and Transformer architecture, enabling comprehensive feature extraction in both spatial and frequency domains. Our encoder employs a hybrid approach combining fast Fourier convolution with central difference convolution, achieving expanded receptive fields while preserving structural details. The decoder features a concurrent interaction architecture comprising a dynamic multi-scale spatial convolution branch and a dual-scale frequency-adjusted Transformer branch, which establish inter-domain communication through cross-attention mechanisms. Notably, the frequency-adjusted Transformer incorporates adaptive frequency allocation that dynamically prioritizes critical frequency components while suppressing non-essential frequency constituents. This mechanism help improve segmentation accuracy and model robustness. Experiments on two benchmark RSI segmentation datasets demonstrate the superior performance of our method, with ablation studies validating the effectiveness of each proposed component.