Abstract <p>Surveillance cameras frequently capture only the rear or side views of criminals, posing a challenge for investigators: how to utilize these suspect images without faces to swiftly locate the face of a suspect in vast surveillance video archives. This study introduces a pioneering YOLOv8-Reidentification + real-time search application (YRRSA) approach, which is a two-stage solution that efficiently processes raw surveillance video and conducts real-time searches for suspect faces. This system combines an enhanced YOLOv8 model with an improved person reidentification (ReID) model. In the first-stage subnetwork, the authors design a multispace-to-depth (MSTD) module and apply it to YOLOv8 to strengthen its ability to extract detailed features. A fusion-based ReID model is designed as the second-stage subnetwork; is composed of a ResNet backbone network and a designed dynamic deformable fusion (DDF) structure. This ReID model can overcome the most significant body and limb deformation issues encountered in person reidentification tasks. Although the performance of these two subnetworks cannot surpass that of the state-of-the-art (SOTA) networks, they have the fewest parameters and the fastest speed among the networks with the same level of performance. The enhanced YOLOv8 model achieves excellent accuracy while maintaining a very small number of parameters. The improved ReID model has a similar accuracy to that of the SOTA approaches but is more than twice as fast as them. This study improves the baseline model to achieve higher accuracy and a faster speed without considering computing resources. Moreover, the unique target matching algorithm (UTMA) is designed to implement facial detection-based criminal matching. Finally, this study conducts actual scenario tests by deploying the proposed system, further highlighting its significance for use in real-world applications.</p> Graphical abstract <p>Criminal investigators often can capture only the backs or profiles of suspects, which contain limited information for case investigations. How to quickly find the frontal image of a suspect within massive surveillance video data by using back images, extract facial information, and compare this information with public security system data is crucial for conducting case investigations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A real-time application employing back- or side-view images of individuals to search for specific faces

  • Cunwei Du,
  • Jiali Deng,
  • Tong Zixuan,
  • Xiaomin Wang,
  • Kai Cheng,
  • Xiaolong Zhou,
  • Ming Liu

摘要

Abstract

Surveillance cameras frequently capture only the rear or side views of criminals, posing a challenge for investigators: how to utilize these suspect images without faces to swiftly locate the face of a suspect in vast surveillance video archives. This study introduces a pioneering YOLOv8-Reidentification + real-time search application (YRRSA) approach, which is a two-stage solution that efficiently processes raw surveillance video and conducts real-time searches for suspect faces. This system combines an enhanced YOLOv8 model with an improved person reidentification (ReID) model. In the first-stage subnetwork, the authors design a multispace-to-depth (MSTD) module and apply it to YOLOv8 to strengthen its ability to extract detailed features. A fusion-based ReID model is designed as the second-stage subnetwork; is composed of a ResNet backbone network and a designed dynamic deformable fusion (DDF) structure. This ReID model can overcome the most significant body and limb deformation issues encountered in person reidentification tasks. Although the performance of these two subnetworks cannot surpass that of the state-of-the-art (SOTA) networks, they have the fewest parameters and the fastest speed among the networks with the same level of performance. The enhanced YOLOv8 model achieves excellent accuracy while maintaining a very small number of parameters. The improved ReID model has a similar accuracy to that of the SOTA approaches but is more than twice as fast as them. This study improves the baseline model to achieve higher accuracy and a faster speed without considering computing resources. Moreover, the unique target matching algorithm (UTMA) is designed to implement facial detection-based criminal matching. Finally, this study conducts actual scenario tests by deploying the proposed system, further highlighting its significance for use in real-world applications.

Graphical abstract

Criminal investigators often can capture only the backs or profiles of suspects, which contain limited information for case investigations. How to quickly find the frontal image of a suspect within massive surveillance video data by using back images, extract facial information, and compare this information with public security system data is crucial for conducting case investigations.