<p>Aerial scene understanding using unmanned aerial vehicles plays a crucial role in disaster response, providing vital information for rescue operations and mitigation strategies. This need has fueled research into disaster scene analysis using deep learning models, particularly for identifying distributed object classes like floodwater, vegetation, and roads. While semantic segmentation is well-suited for analyzing such objects, the creation of manually annotated segmentation datasets is a labor-intensive process. In contrast, bounding box annotations for object detection are faster and easier to generate. Weakly supervised segmentation models can use bounding box labels to generate segmentation masks. However, most existing approaches focus primarily on foreground objects (“things") and overlook critical background elements (“stuff"), such as floodwater and vegetation. This work introduces methods to convert bounding box coordinates of object classes of an image to a semantic segmentation mask. For foreground object segmentation, a straightforward approach using segment anything model (SAM) is used. Background object segmentation is achieved through an initial unsupervised clustering of the image using either K-means clustering or SAM. To classify these clusters accurately, two novel methods are proposed that make use of semantic information derived from bounding boxes: (1) box margin based cluster selection, and (2) bag of pixels based cluster classification. The effectiveness of the framework is validated using the FloodNet dataset. The framework is extended into an end-to-end pipeline by integrating a state-of-the-art object detection model for weakly supervised semantic segmentation. The performance of this proposed pipeline is on par with state-of-the-art supervised segmentation models. Additionally, the best-performing conversion method is applied to the K-Flood dataset to generate ground truth segmentation masks from box annotations with minimal manual effort.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Weakly Supervised Foreground and Background Segmentation Framework for Aerial Flood Image Analysis

  • A. V. Shubhasree,
  • Praveen Sankaran,
  • C. V. Raghu

摘要

Aerial scene understanding using unmanned aerial vehicles plays a crucial role in disaster response, providing vital information for rescue operations and mitigation strategies. This need has fueled research into disaster scene analysis using deep learning models, particularly for identifying distributed object classes like floodwater, vegetation, and roads. While semantic segmentation is well-suited for analyzing such objects, the creation of manually annotated segmentation datasets is a labor-intensive process. In contrast, bounding box annotations for object detection are faster and easier to generate. Weakly supervised segmentation models can use bounding box labels to generate segmentation masks. However, most existing approaches focus primarily on foreground objects (“things") and overlook critical background elements (“stuff"), such as floodwater and vegetation. This work introduces methods to convert bounding box coordinates of object classes of an image to a semantic segmentation mask. For foreground object segmentation, a straightforward approach using segment anything model (SAM) is used. Background object segmentation is achieved through an initial unsupervised clustering of the image using either K-means clustering or SAM. To classify these clusters accurately, two novel methods are proposed that make use of semantic information derived from bounding boxes: (1) box margin based cluster selection, and (2) bag of pixels based cluster classification. The effectiveness of the framework is validated using the FloodNet dataset. The framework is extended into an end-to-end pipeline by integrating a state-of-the-art object detection model for weakly supervised semantic segmentation. The performance of this proposed pipeline is on par with state-of-the-art supervised segmentation models. Additionally, the best-performing conversion method is applied to the K-Flood dataset to generate ground truth segmentation masks from box annotations with minimal manual effort.