Multimodal vision-language models in chest x-ray analysis: a study of generalization, supervision, and robustness
摘要
Multimodal vision-language models (VLMs) are increasingly applied to medical imaging, yet systematic evaluations comparing them with unimodal models across datasets, supervision regimes, and clinical domains remain scarce. Prior studies often focus on a single dataset, specific pathologies, or one supervision setting, leaving unclear how these models generalize under realistic variability. We conduct a systematic evaluation of six leading unimodal and multimodal models for chest X-ray (CXR) classification using four widely adopted datasets: MIMIC-CXR, CheXpert, NIH-14, and PadChest. We assess model behavior in both zero-shot (ZS) and fine-tuned (FT) configurations, with a focus on generalization across pathologies, datasets, and linguistic domains. Our findings show that pretrained multimodal models such as CheXzero and CXR-LLaVA perform strongly in zero-shot scenarios, especially on out-of-distribution data, reflecting their capacity for semantic generalization. However, their performance tends to decline after fine-tuning in cross-lingual or noisy-label contexts, indicating susceptibility to overfitting. In contrast, unimodal models gain substantially from supervised fine-tuning, especially on in-domain data. Limitations include evaluation on seven shared pathologies, CXR imaging only, and use of publicly available pretrained models, which may restrict generalization to other clinical tasks. These findings highlight key trade-offs between generalization, robustness, and adaptability, and suggest promise in hybrid training strategies that integrate multimodal priors with targeted domain supervision.