<p>In this paper, we propose a face recognition system with a deep learning technique. The design uses the separable convolution accelerator with HWCK data scheduling. This scheduling method organizes the weight data according to the number of PEs (Processing Elements), considering hardware resources such as bandwidth and memory size. It is used to accelerate the deep separable convolution model through depthwise convolution, pointwise convolution, and batch normalization. We implement the system on a Xilinx ZCU106 development board, using an SoC architecture with ARM and FPGA to achieve a system-level access control design. The proposed accelerator achieves 222 FPS and 60.8 GOPS on the FaceNet-based network. The power consumption on the Xilinx ZCU106 board is 8.82 W with 6.89 GOPS/W performance. Additionally, our design can retain 94% accuracy on the VGGFACE2 dataset, and 99.2% on the LFW dataset. Compared to previous works, our design demonstrates superior real-time performance and energy efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An SoC-based CNN accelerator for face recognition using HWCK data scheduling

  • Tsung-Han Tsai,
  • Chin-Wei Hsu

摘要

In this paper, we propose a face recognition system with a deep learning technique. The design uses the separable convolution accelerator with HWCK data scheduling. This scheduling method organizes the weight data according to the number of PEs (Processing Elements), considering hardware resources such as bandwidth and memory size. It is used to accelerate the deep separable convolution model through depthwise convolution, pointwise convolution, and batch normalization. We implement the system on a Xilinx ZCU106 development board, using an SoC architecture with ARM and FPGA to achieve a system-level access control design. The proposed accelerator achieves 222 FPS and 60.8 GOPS on the FaceNet-based network. The power consumption on the Xilinx ZCU106 board is 8.82 W with 6.89 GOPS/W performance. Additionally, our design can retain 94% accuracy on the VGGFACE2 dataset, and 99.2% on the LFW dataset. Compared to previous works, our design demonstrates superior real-time performance and energy efficiency.