An efficient NoC-based architecture for convolutional neural networks on multicore systems
摘要
Convolutional neural networks (CNNs) have been extensively applied in many practical applications such as image classification, speech processing, and object recognition. CNNs are growing deeper to obtain higher accuracies. Moreover, the large amount of data exchange between neurons complicates communication. As a result, it is critical to design high-performance processing hardware to handle CNN workloads. The multicore system based on network-on-chip (NoC) architecture is a promising solution for performing CNN workloads because NoC provides high performance and scalable interconnections for on-chip communications. However, the performance of NoC-based design is greatly impacted by matching the topology to the characteristics of CNN workloads since different CNN applications require varying amounts of connection bandwidth. Therefore, we propose a NoC-based architecture for CNN workloads by using long-distance wire links with respect to performance metrics. The performance of the proposed NoC-based design is evaluated for the two CNN models, LeNet and CDBNet models, in terms of network throughput, average latency, and energy consumption. Additionally, an analysis is conducted on the area overhead related to the proposed design. The experimental findings verified that the suggested NoC-based design is a very competitive design in comparison with the other designs for CNN workloads.