In this chapter, we study diverse edge DNN inference serving for multi-user and multi-application at scale, which aims to navigate the three-way trade-off between inference accuracy, latency, and resource cost via jointly optimizing the application configuration adaption, DNN model selection, and edge resource provisioning on-the-fly.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Online Optimization for Diverse Edge DNN Inference Serving at Scale

  • Sen Lin,
  • Zhi Zhou,
  • Zhaofeng Zhang,
  • Xu Chen,
  • Junshan Zhang

摘要

In this chapter, we study diverse edge DNN inference serving for multi-user and multi-application at scale, which aims to navigate the three-way trade-off between inference accuracy, latency, and resource cost via jointly optimizing the application configuration adaption, DNN model selection, and edge resource provisioning on-the-fly.