This paper presents a comparative analysis of building segmentation using pre-trained deep learning models on different data types: RGB images and LiDAR data. The study utilizes the MapAI dataset and evaluates the performance of the Segment Everything Everywhere All at Once (SEEM) model for direct inference and the You Only Look Once (YOLOv8) model fine-tuned for building segmentation. Performance is assessed based on Intersection over Union (IoU) and Boundary Intersection over Union (BIoU) metrics. Our findings reveal that SEEM inference on RGB images yields better results compared to LiDAR data, and a similar trend is observed for the YOLOv8 model. Additionally, in the context of fine-tuning, the YOLOv8 model significantly outperforms the trained shallow model and trained weighted U-Net ensemble on RGB images, demonstrating its superior capability for precise building segmentation. Conversely, for LiDAR images, the customized U-Net model yields better results than YOLOv8, highlighting the importance of model selection based on the data type. This comprehensive evaluation underscores the critical role of model adaptation and data characteristics in remote sensing applications for building extraction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging YOLOv8 Fine-Tuning and SEEM Inference for Accurate Building Segmentation Using RGB and LiDAR Datasets

  • Muhammad Sulaiman,
  • Mina Farmanbar,
  • Ahmed Nabil Belbachir,
  • Chunming Rong

摘要

This paper presents a comparative analysis of building segmentation using pre-trained deep learning models on different data types: RGB images and LiDAR data. The study utilizes the MapAI dataset and evaluates the performance of the Segment Everything Everywhere All at Once (SEEM) model for direct inference and the You Only Look Once (YOLOv8) model fine-tuned for building segmentation. Performance is assessed based on Intersection over Union (IoU) and Boundary Intersection over Union (BIoU) metrics. Our findings reveal that SEEM inference on RGB images yields better results compared to LiDAR data, and a similar trend is observed for the YOLOv8 model. Additionally, in the context of fine-tuning, the YOLOv8 model significantly outperforms the trained shallow model and trained weighted U-Net ensemble on RGB images, demonstrating its superior capability for precise building segmentation. Conversely, for LiDAR images, the customized U-Net model yields better results than YOLOv8, highlighting the importance of model selection based on the data type. This comprehensive evaluation underscores the critical role of model adaptation and data characteristics in remote sensing applications for building extraction.