Human detection from drone-acquired images is an active area of research because of its existence in various domains such as surveillance, search and rescue operations, disaster, and crowd management, among others. This study aims to set a benchmark for human detection in drone-acquired aerial image tasks. It involves evaluating the performance of specific deep learning models capable of performing such tasks and achieving acceptable results. RetinaNet, Faster R-CNN, YOLOv5, YOLOv7, and YOLOv8 are the evaluated deep learning models. The dataset used for the evaluation is the National Taiwan University of Technology (NTUT) 4K drone images dataset produced by the National Taiwan University of Technology. Python, TensorFlow, Keras, and OpenCV are used to develop the deep learning models. The evaluation metrics used for the benchmarking are accuracy, precision, recall, and F1 score to determine which model is the most effective in human detection tasks. The results of the evaluated models show that human detection from drone-acquired images or video is challenging due to the diverse characteristics of the human object and properties of drone images. YOLOv8 demonstrated the highest performance, achieving an average accuracy of 58.75%, precision of 90.12%, recall of 64.34%, and F1 score of 75.08% during testing. The second and third best algorithms are YOLOv7 and YOLOv5, followed by Faster R-CNN, and RetinaNet delivered the lowest performance. The results indicate that the specialized deep learning models for object detection struggled to achieve effective and reliable human detection for most of the complex scenes of the NTUT 4K drone images dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Deep Learning Approaches for Detecting Humans in Drone-Acquired Aerial Images

  • Salama A. Mostafa,
  • Nimal Paran Achuthan,
  • Shihab Hamad Khaleefah,
  • Hawaa A. Obaid,
  • Mohd Farhan Md. Fudzee,
  • Aida Mustapha

摘要

Human detection from drone-acquired images is an active area of research because of its existence in various domains such as surveillance, search and rescue operations, disaster, and crowd management, among others. This study aims to set a benchmark for human detection in drone-acquired aerial image tasks. It involves evaluating the performance of specific deep learning models capable of performing such tasks and achieving acceptable results. RetinaNet, Faster R-CNN, YOLOv5, YOLOv7, and YOLOv8 are the evaluated deep learning models. The dataset used for the evaluation is the National Taiwan University of Technology (NTUT) 4K drone images dataset produced by the National Taiwan University of Technology. Python, TensorFlow, Keras, and OpenCV are used to develop the deep learning models. The evaluation metrics used for the benchmarking are accuracy, precision, recall, and F1 score to determine which model is the most effective in human detection tasks. The results of the evaluated models show that human detection from drone-acquired images or video is challenging due to the diverse characteristics of the human object and properties of drone images. YOLOv8 demonstrated the highest performance, achieving an average accuracy of 58.75%, precision of 90.12%, recall of 64.34%, and F1 score of 75.08% during testing. The second and third best algorithms are YOLOv7 and YOLOv5, followed by Faster R-CNN, and RetinaNet delivered the lowest performance. The results indicate that the specialized deep learning models for object detection struggled to achieve effective and reliable human detection for most of the complex scenes of the NTUT 4K drone images dataset.