Object detection in surveillance camera systems is recognized as a crucial application in the field of computer vision, with numerous uses in security, traffic monitoring, and public management. However, edge devices, such as integrated cameras and compact computing systems, encounter limitations in computational resources. To address this issue, a lightweight Transformer model is proposed, optimized to meet the demand for efficient object detection on edge devices. Specifically, the model is constructed by reducing the number of parameters, utilizing lighter Attention variants, and optimizing memory usage during computation. Experiments conducted on popular object detection datasets such as Pascal VOC show that the model achieves higher accuracy than more complex models like SSD and RetinaNet, while significantly reducing computational costs and meeting real-time requirements on edge devices. Specifically, the model achieves a mean Average Precision (mAP) at IoU 0.5 of 60.4%, with approximately 30M parameters and 16.8 GFlops, demonstrating both accuracy and efficiency. These results confirm the feasibility of using Transformer models in intelligent surveillance systems, paving the way for further advancements in computer vision applications in resource-constrained environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Lightweight Transformer Model for Real-Time Object Detection on Edge Devices in Surveillance Camera Systems

  • Dung Nguyen,
  • Van-Dung Hoang,
  • Van-Tuong-Lan Le

摘要

Object detection in surveillance camera systems is recognized as a crucial application in the field of computer vision, with numerous uses in security, traffic monitoring, and public management. However, edge devices, such as integrated cameras and compact computing systems, encounter limitations in computational resources. To address this issue, a lightweight Transformer model is proposed, optimized to meet the demand for efficient object detection on edge devices. Specifically, the model is constructed by reducing the number of parameters, utilizing lighter Attention variants, and optimizing memory usage during computation. Experiments conducted on popular object detection datasets such as Pascal VOC show that the model achieves higher accuracy than more complex models like SSD and RetinaNet, while significantly reducing computational costs and meeting real-time requirements on edge devices. Specifically, the model achieves a mean Average Precision (mAP) at IoU 0.5 of 60.4%, with approximately 30M parameters and 16.8 GFlops, demonstrating both accuracy and efficiency. These results confirm the feasibility of using Transformer models in intelligent surveillance systems, paving the way for further advancements in computer vision applications in resource-constrained environments.