Performance portability of generated cardiac simulation kernels through automatic dimensioning and load balancing on heterogeneous nodes
摘要
Electrophysiology simulation applications, such as the community-developed openCARP framework for in silico experiments, involve applying a broad range of ionic model kernels with different computational weights and arithmetic intensity characteristics. Efficiently executing these kernels while taking into account variations in kernel execution time and heterogeneous processing unit speeds, is crucial for overall simulation performance. Ensuring that this execution strategy adapts automatically to the underlying hardware architecture is important to ensure performance portability. To address these challenges, this work introduces a method for efficiently distributing ionic model kernels across heterogeneous resources, guided by a resource dimensioning heuristic that adapts to each model’s computational profile. These mechanisms are integrated into openCARP and evaluated on 30 representative ionic models, with a focus on both performance and energy efficiency. We demonstrate that on a node with 8 GPUs, our method achieves a geometric mean speedup of 1.45 compared to using all GPUs, while also improving energy efficiency.