<p>SYCL is a modern royalty-free heterogeneous programming specification maintained by the Khronos Group. Recently, it has become increasingly more prevalent and matured, leading to various assessments of its performance, portability, and programmability. While previous evaluations have mainly focused on X86 CPUs, NVIDIA GPUs, and AMD GPUs, how well SYCL performs on ARM multi-core CPUs is still unknown. In this paper, we evaluate three SYCL implementations (i.e., <Emphasis FontCategory="NonProportional">DPCPP</Emphasis>, <Emphasis FontCategory="NonProportional">AdaptiveCPP</Emphasis>, and <Emphasis FontCategory="NonProportional">MLIR-SYCL</Emphasis>) on ARM multi-core CPUs, to uncover performance traps and offer optimization techniques. We use the <Emphasis FontCategory="NonProportional">SYCL-Bench</Emphasis> benchmark suite to assess the performance of <Emphasis FontCategory="NonProportional">DPCPP</Emphasis>, <Emphasis FontCategory="NonProportional">AdaptiveCPP</Emphasis>, and <Emphasis FontCategory="NonProportional">MLIR-SYCL</Emphasis> against their OpenMP counterparts. We also assess the compiler and runtime overhead to evaluate the usability and productivity of the SYCL implementations. Our empirical results demonstrate that these SYCL implementations can achieve satisfactory performance on ARM multi-core processors. Additionally, we highlight several key optimizations, such as NUMA management, which must be carefully addressed to enhance performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An empirical performance evaluation of SYCL on ARM multi-core processors

  • Hanzheng Liang,
  • Chencheng Deng,
  • Peng Zhang,
  • Jianbin Fang,
  • Tao Tang,
  • Chun Huang

摘要

SYCL is a modern royalty-free heterogeneous programming specification maintained by the Khronos Group. Recently, it has become increasingly more prevalent and matured, leading to various assessments of its performance, portability, and programmability. While previous evaluations have mainly focused on X86 CPUs, NVIDIA GPUs, and AMD GPUs, how well SYCL performs on ARM multi-core CPUs is still unknown. In this paper, we evaluate three SYCL implementations (i.e., DPCPP, AdaptiveCPP, and MLIR-SYCL) on ARM multi-core CPUs, to uncover performance traps and offer optimization techniques. We use the SYCL-Bench benchmark suite to assess the performance of DPCPP, AdaptiveCPP, and MLIR-SYCL against their OpenMP counterparts. We also assess the compiler and runtime overhead to evaluate the usability and productivity of the SYCL implementations. Our empirical results demonstrate that these SYCL implementations can achieve satisfactory performance on ARM multi-core processors. Additionally, we highlight several key optimizations, such as NUMA management, which must be carefully addressed to enhance performance.