Snow’s diverse shapes and sizes pose a significant challenge for single-image desnowing. Current methods typically rely on simple networks to predict the snow masks and subsequently remove snow from images accordingly. However, inaccurate predictions often result in incomplete snow removal. This paper introduces BAT-Net, a novel solution equipped with an image transformer encoder and dual transformer decoders. This architecture enables precise snow mask prediction and comprehensive single image desnowing simultaneously. The image encoder enhances and consolidates features from various layers using Scale Conversion Modules (SCM) and a Feature Aggregation Module (FAM). To ensure accurate snow mask prediction, the snow decoder mirrors the complete transformer structure of the background decoder. Within the background decoder, multiple Bidirectional Attention Modules (BAM) with integrated forward and reverse attention branches effectively leverage snow features and reverse snow features from the snow decoder, facilitating complete snow removal. Moreover, real-world desnowing datasets often contain fallen snow, complicating model training and validation. To address this challenge, we introduce the FallingSnow dataset, exclusively featuring scenes of falling snow. Experimental results across five diverse synthetic and real-world snow removal datasets demonstrate BAT-Net’s significant advancements in addressing the challenges of snow mask prediction and single-image desnowing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Simultaneous Snow Mask Prediction and Single Image Desnowing with a Bidirectional Attention Transformer Network

  • Yongheng Zhang,
  • Danfeng Yan

摘要

Snow’s diverse shapes and sizes pose a significant challenge for single-image desnowing. Current methods typically rely on simple networks to predict the snow masks and subsequently remove snow from images accordingly. However, inaccurate predictions often result in incomplete snow removal. This paper introduces BAT-Net, a novel solution equipped with an image transformer encoder and dual transformer decoders. This architecture enables precise snow mask prediction and comprehensive single image desnowing simultaneously. The image encoder enhances and consolidates features from various layers using Scale Conversion Modules (SCM) and a Feature Aggregation Module (FAM). To ensure accurate snow mask prediction, the snow decoder mirrors the complete transformer structure of the background decoder. Within the background decoder, multiple Bidirectional Attention Modules (BAM) with integrated forward and reverse attention branches effectively leverage snow features and reverse snow features from the snow decoder, facilitating complete snow removal. Moreover, real-world desnowing datasets often contain fallen snow, complicating model training and validation. To address this challenge, we introduce the FallingSnow dataset, exclusively featuring scenes of falling snow. Experimental results across five diverse synthetic and real-world snow removal datasets demonstrate BAT-Net’s significant advancements in addressing the challenges of snow mask prediction and single-image desnowing.