CUDA Core Meets Tensor Core: Characterizing the Performance of Tensor Computation on Modern GPU
摘要
Tensor computation plays a crucial role in diverse domains such as artificial intelligence, engineering simulation, and scientific computing. GPUs, the most commonly used platforms for high-performance computing, have been leveraged to accelerate tensor computation. CUDA Core is a significant component of GPUs, providing parallel computing capacity. Tensor Core, a specialized unit that could perform efficient tensor computation is proposed in recent architectures, Volta and Turing. Although manufacturers announced the performance improvement of Tensor Cores, the work of understanding the characteristics of Tensor Cores and comparing CUDA Cores and Tensor Cores is scarce. In this paper, we characterize the performance of tensor computation on two representative GPUs: Tesla V100 and Tesla T4. Based on experimental results, we provide some insights into using Tensor Cores. Furthermore, our detailed characterization and analysis could help developers in programming tensor-based applications on GPUs.