Tiny Person detection in long-range scenes is a popular and challenging task. Current person detectors have two major issues. Firstly, their performance is poor in the case of tiny and heavily occluded persons. Secondly, they are computation-intensive and have large model sizes, which make them difficult to deploy on resource-limited devices. To solve the above issues, we proposed TPS-YOLO. Based on YOLOv8, we reconstruct the network structure by introducing shallow features of P2 into the feature fusion layers, which helps retain more spatial information important for tiny person detection. We design a fine-grained feature extraction module SPDCA to replace the standard convolution layer in the backbone network to enhance the feature representation of the network. In the feature fusion network, we use a weighted fusion method to fuse multi-scale features, which introduces learnable weights to learn the importance of different input features. We propose a lightweight module named C2f_Efficient, which integrates Depthwise Separable Convolution (DSC) to reduce the model parameters. Furthermore, we apply a model pruning method to further reduce the model’s computational complexity. Experiments on the Tinypersonv2 and VisDrone-person datasets show that TPS-YOLO achieves satisfactory performance in terms of both efficiency and accuracy and has advantages on model lightweight.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TPS-YOLO: The Efficient Tiny Person Detection Network Based on Improved YOLOv8 and Model Pruning

  • Li Yao,
  • Qianni Huang,
  • Yan Wan

摘要

Tiny Person detection in long-range scenes is a popular and challenging task. Current person detectors have two major issues. Firstly, their performance is poor in the case of tiny and heavily occluded persons. Secondly, they are computation-intensive and have large model sizes, which make them difficult to deploy on resource-limited devices. To solve the above issues, we proposed TPS-YOLO. Based on YOLOv8, we reconstruct the network structure by introducing shallow features of P2 into the feature fusion layers, which helps retain more spatial information important for tiny person detection. We design a fine-grained feature extraction module SPDCA to replace the standard convolution layer in the backbone network to enhance the feature representation of the network. In the feature fusion network, we use a weighted fusion method to fuse multi-scale features, which introduces learnable weights to learn the importance of different input features. We propose a lightweight module named C2f_Efficient, which integrates Depthwise Separable Convolution (DSC) to reduce the model parameters. Furthermore, we apply a model pruning method to further reduce the model’s computational complexity. Experiments on the Tinypersonv2 and VisDrone-person datasets show that TPS-YOLO achieves satisfactory performance in terms of both efficiency and accuracy and has advantages on model lightweight.