<p>The rapid evolution of advanced cloud services, together with the significant progress of generative artificial intelligence (GAI), has intensified the demand for image generation systems capable of delivering high perceptual quality, low latency, and flexible user control. Achieving these objectives simultaneously remains challenging, as diffusion models, despite their strong generative capability, require substantial computational resources and iterative inference procedures that are difficult to efficiently deploy in heterogeneous and resource-constrained service environments. To address this challenge, we propose a perception-driven cloud-edge collaborative inference framework tailored for advanced cloud service architectures. The proposed framework integrates an Image Quality Prediction (IQP) module to assess the perceptual sufficiency of intermediate latent states and dynamically guide latency-aware scheduling of the denoising process. This mechanism enables adaptive workload allocation between centralized cloud servers and lightweight edge service nodes, allowing the system to balance computational efficiency and generation fidelity in real time. Furthermore, the framework supports user-driven adjustment of the trade-off between perceptual quality and inference latency, significantly enhancing system flexibility and service adaptability. Extensive experiments on public datasets demonstrate that the proposed approach reduces inference latency by 34.7% and computational cost by 41.2%, while improving SSIM and PSNR by 12.3% and 10.6%, respectively. These results validate the effectiveness and scalability of the proposed framework, highlighting its potential for efficient deployment of diffusion models in modern advanced cloud service environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Perception-driven cloud–edge collaborative inference for efficient deployment of diffusion models in advanced cloud services

  • Yiming Zhang,
  • Donglei Tao,
  • Cheng Wang,
  • Qin Ren,
  • Yi Ru,
  • Ledong An,
  • Yifan Liu

摘要

The rapid evolution of advanced cloud services, together with the significant progress of generative artificial intelligence (GAI), has intensified the demand for image generation systems capable of delivering high perceptual quality, low latency, and flexible user control. Achieving these objectives simultaneously remains challenging, as diffusion models, despite their strong generative capability, require substantial computational resources and iterative inference procedures that are difficult to efficiently deploy in heterogeneous and resource-constrained service environments. To address this challenge, we propose a perception-driven cloud-edge collaborative inference framework tailored for advanced cloud service architectures. The proposed framework integrates an Image Quality Prediction (IQP) module to assess the perceptual sufficiency of intermediate latent states and dynamically guide latency-aware scheduling of the denoising process. This mechanism enables adaptive workload allocation between centralized cloud servers and lightweight edge service nodes, allowing the system to balance computational efficiency and generation fidelity in real time. Furthermore, the framework supports user-driven adjustment of the trade-off between perceptual quality and inference latency, significantly enhancing system flexibility and service adaptability. Extensive experiments on public datasets demonstrate that the proposed approach reduces inference latency by 34.7% and computational cost by 41.2%, while improving SSIM and PSNR by 12.3% and 10.6%, respectively. These results validate the effectiveness and scalability of the proposed framework, highlighting its potential for efficient deployment of diffusion models in modern advanced cloud service environments.