<p>Semantic segmentation of building acoustic materials is crucial for applications in Building Information Modeling (BIM) and digital twin construction, a task challenged by the high computational demands and architectural limitations of general-purpose models. To address this, this paper proposes a computationally efficient and accurate segmentation framework, named OptimizedSwinUNet. Our proposed architecture enhances the SwinUNet baseline through three targeted optimizations designed to improve both accuracy and efficiency: (1) a hybrid window attention (HWA) module that fuses local and global features with minimal overhead; (2) an integrated Squeeze-and-Excitation (SE) module for channel-wise feature refinement; and (3) an optimized sampling strategy using depthwise separable convolutions that significantly reduces the model’s parameter count. On our newly constructed&#xa0;acoustic material segmentation (AMS) dataset—a challenging benchmark comprising 1869 high-resolution images of five material categories, captured across diverse real-world construction sites and characterized by significant class imbalance and visually ambiguous boundaries—comprehensive experiments demonstrate the superiority of our framework. Compared to the baseline SwinUNet,&#xa0;our OptimizedSwinUNet achieves a mean Intersection over Union (IoU) of 89.5%, representing a notable improvement, while simultaneously reducing the parameter count by 1.9%.&#xa0;Furthermore, the utility of these high-fidelity segmentation results is demonstrated through their integration into a 3D reconstruction pipeline, where they lead to a&#xa0;33.3% reduction in the geometric root mean square error (RMSE)&#xa0;of the final model. This work presents an effective and resource-aware framework for automated material analysis, showcasing strong potential for engineering applications. Validation on the public synapse dataset further confirms our model’s strong generalization ability, outperforming the baseline by a large margin.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

OptimizedSwinUNet: a computationally efficient framework for joint semantic segmentation and 3D modeling of acoustic materials

  • Zheng Ting,
  • Pang Jinxiang,
  • Liu Xuewen,
  • Zhang Zhenshan

摘要

Semantic segmentation of building acoustic materials is crucial for applications in Building Information Modeling (BIM) and digital twin construction, a task challenged by the high computational demands and architectural limitations of general-purpose models. To address this, this paper proposes a computationally efficient and accurate segmentation framework, named OptimizedSwinUNet. Our proposed architecture enhances the SwinUNet baseline through three targeted optimizations designed to improve both accuracy and efficiency: (1) a hybrid window attention (HWA) module that fuses local and global features with minimal overhead; (2) an integrated Squeeze-and-Excitation (SE) module for channel-wise feature refinement; and (3) an optimized sampling strategy using depthwise separable convolutions that significantly reduces the model’s parameter count. On our newly constructed acoustic material segmentation (AMS) dataset—a challenging benchmark comprising 1869 high-resolution images of five material categories, captured across diverse real-world construction sites and characterized by significant class imbalance and visually ambiguous boundaries—comprehensive experiments demonstrate the superiority of our framework. Compared to the baseline SwinUNet, our OptimizedSwinUNet achieves a mean Intersection over Union (IoU) of 89.5%, representing a notable improvement, while simultaneously reducing the parameter count by 1.9%. Furthermore, the utility of these high-fidelity segmentation results is demonstrated through their integration into a 3D reconstruction pipeline, where they lead to a 33.3% reduction in the geometric root mean square error (RMSE) of the final model. This work presents an effective and resource-aware framework for automated material analysis, showcasing strong potential for engineering applications. Validation on the public synapse dataset further confirms our model’s strong generalization ability, outperforming the baseline by a large margin.