Optimization Strategies and Algorithms for Accelerating CNN on FPGA: A Comprehensive Review
摘要
Recent advancements in signal and image processing algorithms have been widely applied in various sectors, including medical, industrial, robotics, autonomous vehicles, etc. Despite these advancements, achieving optimal processing speed remains a significant challenge. With this insight, Convolutional Neural Network (CNN) has remarkably exploited their power of intelligence in diverse areas. However, their high computational demands and substantial memory bandwidth requirements often constrain their performance on conventional CPUs. This underscores the importance of hardware accelerators, such as Generic Processing Unit (GPU), Field Programmable Gate Array (FPGA), and Application Specific Integrated Circuit (ASIC). Among these accelerators, FPGA, with their reconfigurable architecture and inherent parallelism, offer significant advantages over GPU and ASIC for CNN acceleration, providing flexibility and efficiency in adapting to diverse models and tasks. Furthermore, FPGAs exhibit optimal energy efficiency, making them an attractive choice for CNN acceleration in resource-constrained environments. This comprehensive review provides an in-depth analysis of CNN accelerators implemented on FPGA, exploring architectures, acceleration strategies, and optimization challenges, providing valuable insights for researchers involved in hardware implementation of CNN models.