Edge-Cloud Cooperative Inference for Semantic Segmentation Models
摘要
In the era of pervasive computing, cloud-edge collaboration has become a crucial strategy for enhancing the efficiency and responsiveness of computational tasks. This paper explores advanced inference methods for segmentation models within a cloud-edge collaborative framework. We investigate paradigms that leverage the computational power of cloud servers and the proximity of edge devices to deliver high-precision and low-latency segmentation results. Our study introduces a novel hybrid inference architecture that dynamically partitions the encoder layers of the segmentation model between cloud and edge resources based on real-time computational and network conditions. Additionally, we compress images into feature to reduce transmission latency. Through extensive experiments, we demonstrate that our approach significantly reduces response times and computational overhead while maintaining segmentation accuracy. Specifically, the segmentation performance, measured in terms of Mean Intersection over Union (MIoU), decreases by only 1 point. In contrast, inference speed sees a notable improvement of 30.3%.