The Locally Optimal Block Preconditioned Conjugate Gradient (LOBPCG) algorithm is an effective algorithm for solving large-scale sparse eigenvalue problems in various scientific and engineering applications. Based on a distributed CPU-GPU heterogeneous environment, this paper presents an optimized approach for the efficient solution of sparse eigenvalue problems utilizing the LOBPCG algorithm. To conceal intra-node data transfer and reduce the influence of the cross-node MPI communication overhead, we employ a sparse matrix regular two-dimensional partitioning scheme. To enhance GPU data throughput and data reuse, we consolidate multiple sparse matrix-vector multiplications (SPMV) into a single sparse matrix-matrix multiplication (SPMM). Furthermore, for the irregular dense matrix-matrix multiplication (GEMM) involved in the LOBPCG algorithm, we design an adaptive batched GEMM technique for GPU execution. Extensive experimental results demonstrate the computational efficiency and scalability of our approach. Our implementation achieves a significant speedup in solving typical matrices compared to the SLEPc library. The strong scalability performance exceeds 58%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Implementation of the LOBPCG Algorithm on a CPU-GPU Cluster

  • Yang Liu,
  • Yonghua Zhao,
  • Zexin Wang,
  • Rongfeng Huang,
  • Dingye Zhang,
  • Xinyin Zhang

摘要

The Locally Optimal Block Preconditioned Conjugate Gradient (LOBPCG) algorithm is an effective algorithm for solving large-scale sparse eigenvalue problems in various scientific and engineering applications. Based on a distributed CPU-GPU heterogeneous environment, this paper presents an optimized approach for the efficient solution of sparse eigenvalue problems utilizing the LOBPCG algorithm. To conceal intra-node data transfer and reduce the influence of the cross-node MPI communication overhead, we employ a sparse matrix regular two-dimensional partitioning scheme. To enhance GPU data throughput and data reuse, we consolidate multiple sparse matrix-vector multiplications (SPMV) into a single sparse matrix-matrix multiplication (SPMM). Furthermore, for the irregular dense matrix-matrix multiplication (GEMM) involved in the LOBPCG algorithm, we design an adaptive batched GEMM technique for GPU execution. Extensive experimental results demonstrate the computational efficiency and scalability of our approach. Our implementation achieves a significant speedup in solving typical matrices compared to the SLEPc library. The strong scalability performance exceeds 58%.