Deep Learning-Based Pertinent Video Frame Detection to Compressed Surveillance Video
摘要
The proliferation of closed-circuit television (CCTV) cameras has surged significantly, as they are used for security, safety, and monitoring purposes across diverse sectors. However, as the demand for CCTV continues to rise, the major issue associated with surveillance systems is surveillance video storage capacity. The CCTV videos are stored on a local drive or cloud which has limited storage capacity. As the use of surveillance cameras increases, it produce vast amounts of data, which quickly overwhelm local hard drives or cloud storage repositories. Due to this, surveillance video were deleted after a certain period which escalates the risk of loss of potentially valuable information. To tackle this issue, we proposed a deep learning-based pertinent frame detection and compression (D&C) model. The proposed D&C model consists of three phases: (i) data engineering, (ii) pertinent frame detection module and (iii) similarity identification module. We used ten surveillance videos of automated trailer machines (ATM) under two different scenarios to test our proposed D&C model. Experimental results show that in the pertinent frame detection model YOLOv9 surpasses YOLOv5, YOLOv7 and YOLOv8 in terms of speed and accuracy. Using the proposed D&C model, we achieved the highest compression of 98.96% and 68.64% lowest by maintaining the same resolution and frame per second (FPS).