ViTAttUNet+: an architecture for accurate building footprint delineation in aerial imagery
摘要
A substantial volume of remote sensing data is acquired daily, thereby enabling advanced computer vision applications in urban planning, disaster management, and environmental monitoring. Building segmentation from high‑resolution aerial imagery is rendered challenging by diverse structures, intricate backgrounds, and large‑scale variations—limitations that are encountered by traditional CNNs. In the present study, ViTAttUNet+ is introduced as a novel hybrid model in which convolutional neural networks and vision transformers (ViTs) are synergistically integrated with attention gates, atrous spatial pyramid pooling (ASPP), and squeeze‑and‑excitation (SE) blocks so that both local and global features can be captured using compact 256 × 256 inputs. The core contributions are presented as follows: (1) a hybrid encoder–decoder architecture is proposed that balances fine‑grained detail with long‑range context; (2) the targeted integration of attention gates, ASPP, and SE blocks is employed for adaptive multi‑scale feature fusion; and (3) a comprehensive evaluation is performed on the AIRS high‑resolution aerial imagery dataset. It is demonstrated by experiments that a mean IoU of 0.879 and a Dice score of 0.934 are achieved by ViTAttUNet+, corresponding to improvements of + 3.8% and + 2.2%, respectively, over state‑of‑the‑art baselines. These results are indicative of the superiority of the proposed approach for the precise extraction of building footprints.