Deploying Large Language Model on Cloud-Edge Architectures: A Case Study for Conversational Historical Characters
摘要
This work analyzes the deployment of conversational agents based on large language models (LLMs) in cloud-edge architectures, placing emphasis on scalability, efficiency and real-time performance. Through a case study, we present a web application that allows users to interact with an augmented reality avatar that impersonates a historical character. The agent, powered by an LLM delivers immersive and contextually coherent dialogues. We discuss the solutions adopted to manage latency and distribute the computational load between the cloud, which takes care of language processing, and the edge nodes, ensuring a smooth user experience. The results obtained demonstrate how accurate design can optimize the use of LLMs in distributed environments, offering advanced and high-performance interactions even in applications with high reactivity and customization requirements.