Dual sparsity aware PE networks for CNN accelerators in edge AI deployments
摘要
This work proposes a dual-sparsity-aware processing element (PE) incorporated within the systolic array network, which can avoid the ineffective computations resulting from the activation and weight parameters in computationally expensive convolution operations. The bus-specific clock gating is implemented across the registers associated with the adder units in MAC modules to minimize unnecessary signal transitions. This approach utilizes dual-port memory partitioning and memory splitting techniques to enhance memory access efficiency and improve parallelism. The results show that our method achieves a 1.58