Secure Elasticsearch Clusters on HPC Systems for Sensitive Data
摘要
Data catalogs are an established tool to integrate heterogeneous data, enrich raw data with semantic meaningful metadata, and make data easily searchable, maintainable, and shareable. This helps, for instance, data scientists to manage large data sets, which are often required for state-of-the-art artificial intelligence research. Driven by the increasing computing demand of these data-intensive projects, High-performance Computing (HPC) providers have to address the specific demands of these projects to attract this new user group to HPC systems. One particularly challenging domain are the life sciences working with highly-regulated, sensitive health data. This paper presents a workflow to deploy on-demand Elasticsearch (ES) clusters in user space on HPC systems, providing a backend for direct usage or user-defined, higher-order data catalog functionalities. Therewith, it augments the capabilities of a parallel file system by allowing processing of user-defined metadata. Two different encryption techniques are presented and used in two different use cases to systematically benchmark the developed setup, show its general scalability, and highlight important considerations when adapting it to a new use case. It is shown, that scaling out ES clusters has to be done using thorough data and workload modeling since larger clusters can be either beneficial or harmful to different workloads.