DLUSEdge: Dynamic Load–Unload Scheduling for Localized LLMs on Resource-Constrained Edge
摘要
Deploying large language models (LLMs) on resource-constrained edge device presents significant challenges due to their high computational and memory demands. This paper introduces DLUSEdge, an efficient algorithmic framework designed to dynamically manage the loading and unloading of quantized LLMs on edge devices. The framework employs time-bound scheduling to optimize task execution while minimizing resource overhead. Four quantized LLMs, including qwen2.5:0.5b-instruct and granite3-moe:1b-instruct-q4_K_M, were evaluated in real-world scenarios, demonstrating task latency as low as