Resource Management for GPT-Based Model Deployed on Clouds: Challenges, Solutions, and Future Directions
摘要
The widespread adoption of large language models (LLMs), such as the Generative Pre-trained Transformer (GPT), on cloud computing platforms (e.g., Azure) has resulted in a substantial increase in resource demand. This increase poses significant challenges for resource management within cloud environments. This paper aims to highlight these challenges by initially delineating the unique characteristics of resource management for GPT-based models. Subsequently, we analyze the specific challenges faced by resource management when applied to GPT-based models deployed on cloud platforms, and we propose algorithms for resource profiling and prediction concerning inference requests. To facilitate effective resource management, we present a comprehensive resource management framework that includes resource profiling and forecasting methodologies specifically designed for GPT-based models. Additionally, we discuss the future directions for resource management in the context of GPT-based models, emphasizing potential areas for further exploration and improvement. Through this analysis, we aim to provide valuable insights into resource management for GPT-based models deployed in cloud environments.