Medical image analysis faces challenges regarding data privacy and heterogeneity, especially in federated learning contexts, where efficient and accurate image classification remains difficult. Prompt learning in pre-trained vision-language models has shown strong adaptability across various tasks. Recent studies have integrated robust pre-trained models into federated learning systems to reduce communication costs and enable local training in data-limited environments. This approach offers significant advantages for medical image processing with limited computational resources. While previous research has shown promising results in federated natural image analysis, its application to medical imaging remains limited. We propose FedEM-AC, an enhanced federated learning approach, integrating the EM algorithm and the CLIP model with a multi-layer visual encoding attention fusion mechanism to improve medical image classification accuracy. Our EM-based prompt aggregation strategy extracts common semantics from prompt vectors, overcoming the limitations of traditional sample-size weighted averaging. Additionally, we propose a module that enhances the capability of Optimal Transport in aligning images with text by integrating attention matrices from CLIP. Comprehensive experiments in federated learning contexts demonstrate that our approach outperforms other leading methods in medical image classification, excelling at lesion targeting and convergence speed.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FedEM-AC: Enhancing Federated Learning for Medical Image Classification Through EM Aggregation and Adaptive Multi-Layer Attention Fusion

  • Shuo Dai,
  • Fei Dai,
  • Huan Yang,
  • Songsen Yu

摘要

Medical image analysis faces challenges regarding data privacy and heterogeneity, especially in federated learning contexts, where efficient and accurate image classification remains difficult. Prompt learning in pre-trained vision-language models has shown strong adaptability across various tasks. Recent studies have integrated robust pre-trained models into federated learning systems to reduce communication costs and enable local training in data-limited environments. This approach offers significant advantages for medical image processing with limited computational resources. While previous research has shown promising results in federated natural image analysis, its application to medical imaging remains limited. We propose FedEM-AC, an enhanced federated learning approach, integrating the EM algorithm and the CLIP model with a multi-layer visual encoding attention fusion mechanism to improve medical image classification accuracy. Our EM-based prompt aggregation strategy extracts common semantics from prompt vectors, overcoming the limitations of traditional sample-size weighted averaging. Additionally, we propose a module that enhances the capability of Optimal Transport in aligning images with text by integrating attention matrices from CLIP. Comprehensive experiments in federated learning contexts demonstrate that our approach outperforms other leading methods in medical image classification, excelling at lesion targeting and convergence speed.