<p>Recent advancements in large language models (LLMs) have revolutionized artificial intelligence, yet their computational demands pose significant challenges for efficient deployment. A major issue is handling diverse query responses efficiently, which motivates the need to predict response lengths and optimize the batch processes. In this paper, we thoroughly analyze the challenges involved in the LLM response length prediction task and propose a new framework that treats it as an uncertainty-aware regression problem. We benchmark four uncertainty quantification methods, including both Frequentist and Bayesian approaches, and find that evidential deep learning (EDL) is the most effective and efficient for this task. Furthermore, our case study demonstrates that our approach averagely reduces inference time by 38.14% and 20.50% compared with random batching and the state-of-the-art method, respectively, showcasing the potential of uncertainty-aware response length predictions in optimizing LLM inference.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncertainty-aware large language model response length perception

  • Bin Shi,
  • Bo Dong,
  • Qinghua Zheng

摘要

Recent advancements in large language models (LLMs) have revolutionized artificial intelligence, yet their computational demands pose significant challenges for efficient deployment. A major issue is handling diverse query responses efficiently, which motivates the need to predict response lengths and optimize the batch processes. In this paper, we thoroughly analyze the challenges involved in the LLM response length prediction task and propose a new framework that treats it as an uncertainty-aware regression problem. We benchmark four uncertainty quantification methods, including both Frequentist and Bayesian approaches, and find that evidential deep learning (EDL) is the most effective and efficient for this task. Furthermore, our case study demonstrates that our approach averagely reduces inference time by 38.14% and 20.50% compared with random batching and the state-of-the-art method, respectively, showcasing the potential of uncertainty-aware response length predictions in optimizing LLM inference.