<p>Heterogeneous computing and exploiting integrated CPU–GPU architectures has become a&#xa0;clear current trend since the flattening of Moore’s Law. In this work, we propose a&#xa0;numerical and algorithmic re-design of a&#xa0;p-adaptive quadrature-free discontinuous Galerkin (DG) method for the shallow water equations. Our new approach separates the computations of the non-adaptive (lower-order) and adaptive (higher-order) parts of the discretization from each other. Thereby, we can overlap computations of the lower-order and the higher-order DG solution components. Furthermore, we investigate execution times of main computational kernels and use automatic code generation to optimize their distribution between the CPU and GPU. Several setups, including a&#xa0;prototype of a&#xa0;tsunami simulation in a&#xa0;tide-driven flow scenario, are investigated, and the results show that significant performance improvements can be achieved in suitable setups.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

p-adaptive discontinuous Galerkin method for the shallow water equations on heterogeneous computing architectures

  • Sara Faghih-Naini,
  • Vadym Aizinger,
  • Sebastian Kuckuk,
  • Richard Angersbach,
  • Harald Köstler

摘要

Heterogeneous computing and exploiting integrated CPU–GPU architectures has become a clear current trend since the flattening of Moore’s Law. In this work, we propose a numerical and algorithmic re-design of a p-adaptive quadrature-free discontinuous Galerkin (DG) method for the shallow water equations. Our new approach separates the computations of the non-adaptive (lower-order) and adaptive (higher-order) parts of the discretization from each other. Thereby, we can overlap computations of the lower-order and the higher-order DG solution components. Furthermore, we investigate execution times of main computational kernels and use automatic code generation to optimize their distribution between the CPU and GPU. Several setups, including a prototype of a tsunami simulation in a tide-driven flow scenario, are investigated, and the results show that significant performance improvements can be achieved in suitable setups.