Privacy-Preserving Chest X-Ray Report Generation via Multimodal Federated Learning with ViT and GPT2
摘要
The automated generation of radiology reports from chest X-ray images holds significant promise for enhancing diagnostic workflows while preserving patient privacy. Traditional centralized approaches often require sensitive data transfer, raising privacy concerns. To address this, the study proposes a Multimodal Federated Learning (FL) framework for chest X-ray report generation using the IU-X-ray dataset. The system utilizes a Vision Transformer (ViT) as the encoder and GPT2 as the report generator, enabling decentralized training without sharing raw data. Three FL aggregation strategies—FedAvg, Krum Aggregation, and a novel Loss-aware Federated Averaging (L-FedAvg)—were evaluated. Among these, Krum Aggregation showed superior performance across lexical and semantic evaluation metrics such as ROUGE, BLEU, BERTScore, and RaTEScore. To enhance interpretability, attention maps were used to explain the generated reports. The results indicate that FL can match or surpass centralized models in generating clinically relevant and semantically rich reports. This lightweight, privacy-preserving framework paves the way for collaborative medical AI development without compromising data confidentiality.