<p>In response to issues such as incomplete contour segmentation, blurred boundaries, and small-building misclassification in remote sensing images, this paper proposes an MR-DeepLabv3+ network. The network integrates MixConv (dataset-adapted multi-scale convolutional kernels: 3 ×  3/5 × 5/7 × 7) to enhance multi-scale feature capture and a segmentation-optimized R-Drop Loss (decoder-level channel-wise masking with dynamic KL divergence weight) to reinforce noise robustness. Experiments are conducted on three distinct building datasets: Self-building (1270 images, 10–50 pixel slender buildings), WHU (8170 images, 0.075&#xa0;m resolution dense small buildings), and Massachusetts (151 images, 340 km<sup>2</sup> large urban clusters). The experimental results show that the MR-DeepLabv3 + achieves Acc, MIoU, and FWIoU of 98.34%, 88.93%, 96.88% (Self-building), 98.22%, 88.56%, 97.18% (WHU), and 88.32%, 79.33%, 86.53% (Massachusetts), outperforming the baseline DeepLabv3+ and 4 recent transformer models. MR-DeepLabv3 + balances model compactness and inference efficiency, making it well-suited for remote sensing image segmentation tasks, especially in scenarios with computational resource constraints. Ultimately, it is proven that the method effectively improves building segmentation accuracy and addresses small-building missing issues, with practical value for UAV-based mapping.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Precise building semantic segmentation in remote sensing images via MR-DeepLabv3+ network

  • Yiming Wang,
  • Lunhua Shang,
  • Yu Liu,
  • Zeyu Sun,
  • Yanyang Lu,
  • Xunwang Shi,
  • Tianhao Wang,
  • Yifan Liu

摘要

In response to issues such as incomplete contour segmentation, blurred boundaries, and small-building misclassification in remote sensing images, this paper proposes an MR-DeepLabv3+ network. The network integrates MixConv (dataset-adapted multi-scale convolutional kernels: 3 ×  3/5 × 5/7 × 7) to enhance multi-scale feature capture and a segmentation-optimized R-Drop Loss (decoder-level channel-wise masking with dynamic KL divergence weight) to reinforce noise robustness. Experiments are conducted on three distinct building datasets: Self-building (1270 images, 10–50 pixel slender buildings), WHU (8170 images, 0.075 m resolution dense small buildings), and Massachusetts (151 images, 340 km2 large urban clusters). The experimental results show that the MR-DeepLabv3 + achieves Acc, MIoU, and FWIoU of 98.34%, 88.93%, 96.88% (Self-building), 98.22%, 88.56%, 97.18% (WHU), and 88.32%, 79.33%, 86.53% (Massachusetts), outperforming the baseline DeepLabv3+ and 4 recent transformer models. MR-DeepLabv3 + balances model compactness and inference efficiency, making it well-suited for remote sensing image segmentation tasks, especially in scenarios with computational resource constraints. Ultimately, it is proven that the method effectively improves building segmentation accuracy and addresses small-building missing issues, with practical value for UAV-based mapping.