<p>Image segmentation remains a pivotal challenge in computer vision, particularly in complex scenarios requiring fine-grained feature discrimination. Current approaches often suffer from inefficient feature utilization and local detail loss during semantic segmentation. To address these limitations, we propose a novel deep neural network with multi-scale attention fusion for accurate fine-grained image segmentation and the lightweight architecture ensures computational efficiency without sacrificing accuracy. Our approach integrates three key components: the Dynamic Spatial-Atrous Spatial Pyramid Pooling (DSA-ASPP) module, which combines depthwise separable convolution with adaptive dilation rates to reduce parameters; a multi-scale attention fusion mechanism which hierarchically integrates features to enhance local texture discriminability and minimizing computational overhead. and the PreactResNet-ECA, a pre-activated residual network with channel-wise attention optimized for fine-grained feature interaction. Experimental results on CamVid and Cityscapes datasets demonstrate the superior performance of our proposed model, achieving mean intersection-over-union (mIoU) scores of 69.6% and 73.6%, respectively, with inference speeds reaching 255.8 FPS. Furthermore, evaluations on fine-grained datasets (CUB-200-2011 and Stanford Dogs) reveal that our PreactResNet-based model outperforms state-of-the-art approaches, attaining accuracies of 93.0% and 97.0%. The framework effectively preserves local texture details, reduces pixel-level misclassification, and offers a balanced trade-off between accuracy and computational efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A lightweight deep neural network with attention fusion for fine-grained image segmentation in complex scenes

  • Pingshan Liu,
  • Jiangli Liu,
  • Jinfeng Li,
  • Guangyan Huang

摘要

Image segmentation remains a pivotal challenge in computer vision, particularly in complex scenarios requiring fine-grained feature discrimination. Current approaches often suffer from inefficient feature utilization and local detail loss during semantic segmentation. To address these limitations, we propose a novel deep neural network with multi-scale attention fusion for accurate fine-grained image segmentation and the lightweight architecture ensures computational efficiency without sacrificing accuracy. Our approach integrates three key components: the Dynamic Spatial-Atrous Spatial Pyramid Pooling (DSA-ASPP) module, which combines depthwise separable convolution with adaptive dilation rates to reduce parameters; a multi-scale attention fusion mechanism which hierarchically integrates features to enhance local texture discriminability and minimizing computational overhead. and the PreactResNet-ECA, a pre-activated residual network with channel-wise attention optimized for fine-grained feature interaction. Experimental results on CamVid and Cityscapes datasets demonstrate the superior performance of our proposed model, achieving mean intersection-over-union (mIoU) scores of 69.6% and 73.6%, respectively, with inference speeds reaching 255.8 FPS. Furthermore, evaluations on fine-grained datasets (CUB-200-2011 and Stanford Dogs) reveal that our PreactResNet-based model outperforms state-of-the-art approaches, attaining accuracies of 93.0% and 97.0%. The framework effectively preserves local texture details, reduces pixel-level misclassification, and offers a balanced trade-off between accuracy and computational efficiency.