Establishing Academic AI Lab Infrastructures: Challenges and Lessons
摘要
The recent advances in Artificial Intelligence (AI) within the past decade were mainly fueled by developments in high-performance computingHigh-performance computing (HPC). AI and deep learning have revolutionized numerous disciplines. High-performance computingHigh-performance computing (HPC) provides the necessary infrastructure to meet the computational demands of AI, allowing researchers to explore complex architectures and large-scale datasets. This paper aims to present a blueprint for the design of the infrastructure that can handle such large AI models at a medium-sized research facility level, including the hardware and software stacks. With the latest advances in open-source HPCHigh-performance computing software and hardware, it is now possible for small to medium-sized organizations to build computing clusters using off-the-shelf hardware and software components. We present a case study in designing and implementing an advanced AI computing cluster within a GCC-based academic institution using open-source building blocks (software and hardware) to serve both research and training in the AI field. The paper outlines a framework adopted for the design of the HPCHigh-performance computing cluster as well as the challenges and remedies adopted, including best-practice cases adopted from collaboration with major international open computing centers (e.g., CERN HPC Center in Geneva, etc.) for building the infrastructure. The design principles for the lab include calculation of needed scalability, handling of big data, and cluster size. The model framework identifies the weight of each part of the design process, with phases including design, operation and maintenance, training and knowledge transfer, and other cross-cutting factors within all phases. Finally, challenges are identified and outlined.