<p>The automated generation of radiology reports from chest X-ray images holds significant promise for enhancing diagnostic workflows while preserving patient privacy. Traditional centralized approaches often require sensitive data transfer, raising privacy concerns. To address this, the study proposes a Multimodal Federated Learning (FL) framework for chest X-ray report generation using the IU-X-ray dataset. The system utilizes a Vision Transformer (ViT) as the encoder and GPT2 as the report generator, enabling decentralized training without sharing raw data. Three FL aggregation strategies—FedAvg, Krum Aggregation, and a novel Loss-aware Federated Averaging (L-FedAvg)—were evaluated. Among these, Krum Aggregation showed superior performance across lexical and semantic evaluation metrics such as ROUGE, BLEU, BERTScore, and RaTEScore. To enhance interpretability, attention maps were used to explain the generated reports. The results indicate that FL can match or surpass centralized models in generating clinically relevant and semantically rich reports. This lightweight, privacy-preserving framework paves the way for collaborative medical AI development without compromising data confidentiality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Privacy-Preserving Chest X-Ray Report Generation via Multimodal Federated Learning with ViT and GPT2

  • Md. Zahid Hossain,
  • Mustofa Ahmed,
  • Most. Sharmin Sultana Samu,
  • Md. Rakibul Islam

摘要

The automated generation of radiology reports from chest X-ray images holds significant promise for enhancing diagnostic workflows while preserving patient privacy. Traditional centralized approaches often require sensitive data transfer, raising privacy concerns. To address this, the study proposes a Multimodal Federated Learning (FL) framework for chest X-ray report generation using the IU-X-ray dataset. The system utilizes a Vision Transformer (ViT) as the encoder and GPT2 as the report generator, enabling decentralized training without sharing raw data. Three FL aggregation strategies—FedAvg, Krum Aggregation, and a novel Loss-aware Federated Averaging (L-FedAvg)—were evaluated. Among these, Krum Aggregation showed superior performance across lexical and semantic evaluation metrics such as ROUGE, BLEU, BERTScore, and RaTEScore. To enhance interpretability, attention maps were used to explain the generated reports. The results indicate that FL can match or surpass centralized models in generating clinically relevant and semantically rich reports. This lightweight, privacy-preserving framework paves the way for collaborative medical AI development without compromising data confidentiality.