Enhanced Building Extraction via STMC-UNet: Integrating Super Token Transformer and Multi-scale Convolution
摘要
Accurate extraction of buildings from high-resolution remote sensing images is pivotal for urban planning and sustainable development. While existing CNN- and Transformer-based semantic segmentation methods have advanced, they often struggle with long-range dependencies and high computational complexity, respectively. We propose STMC-UNet, a hybrid network combining the Super Token Vision Transformer with multi-scale convolution, to address these challenges. Super Token Vision Transformer reduces computational complexity while enhancing global context modeling through a super token self-attention mechanism. The dual-branch multi-scale feature extraction module further improves local detail extraction efficiency. Additionally, a full-scale feature refinement skip connection and a hybrid feature fusion module are designed to suppress redundant information and achieve refined feature fusion. Experiments on the Massachusetts, WHU, and Inria datasets demonstrate that STMC-UNet achieves IoU scores of 73.38%, 91.15%, and 83.83%, respectively, showcasing its practical value in high-precision building extraction tasks.