In the rapid advancement of artificial intelligence, the acceleration of convolutional neural networks (CNNs) for network security has become a critical area of research, especially for applications requiring real-time processing such as intrusion detection, malware analysis, and secure communications. This paper explores the use of Data Processing Unit (DPU) incorporating Field Programmable Gate Array (FPGA) to implement CNN accelerators for network security tasks. DPU, particularly those utilizing FPGA, are considered a promising hardware platform due to their flexibility and efficiency. We have designed and implemented several architectures, including optimized line buffer structures and pulsating array structures ported from Google’s TPU design, to achieve pipelined convolution operations and acceleration. We compared three different accelerator architectures: a dual convolution kernel parallel pulsating array, a dual convolution kernel parallel addition-tree structure, and an enhanced four convolution kernel parallel pulsating array. The findings show that the enhanced four convolution kernel parallel pulsating array architecture performs ten times faster than traditional CPU while maintaining low resource consumption. The other two architectures also achieved acceleration to varying degrees. These results confirm the feasibility of DPU-based CNN accelerators for network security and emphasize the important balance between computational speed and resource efficiency in designing such systems. This paper points out directions for future research to optimize resource utilization and expand application scenarios, providing valuable insights into the field of AI acceleration and paving the way for the development of high-performance, energy-efficient AI applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inference Acceleration Method for Embedded Data Processing Units Based on Systolic Arrays

  • Yuting Li,
  • Haoyang Bai,
  • Ming Jin,
  • Tiangao Piao,
  • Han Wang,
  • Yitao Xiao

摘要

In the rapid advancement of artificial intelligence, the acceleration of convolutional neural networks (CNNs) for network security has become a critical area of research, especially for applications requiring real-time processing such as intrusion detection, malware analysis, and secure communications. This paper explores the use of Data Processing Unit (DPU) incorporating Field Programmable Gate Array (FPGA) to implement CNN accelerators for network security tasks. DPU, particularly those utilizing FPGA, are considered a promising hardware platform due to their flexibility and efficiency. We have designed and implemented several architectures, including optimized line buffer structures and pulsating array structures ported from Google’s TPU design, to achieve pipelined convolution operations and acceleration. We compared three different accelerator architectures: a dual convolution kernel parallel pulsating array, a dual convolution kernel parallel addition-tree structure, and an enhanced four convolution kernel parallel pulsating array. The findings show that the enhanced four convolution kernel parallel pulsating array architecture performs ten times faster than traditional CPU while maintaining low resource consumption. The other two architectures also achieved acceleration to varying degrees. These results confirm the feasibility of DPU-based CNN accelerators for network security and emphasize the important balance between computational speed and resource efficiency in designing such systems. This paper points out directions for future research to optimize resource utilization and expand application scenarios, providing valuable insights into the field of AI acceleration and paving the way for the development of high-performance, energy-efficient AI applications.