Intersectional fairness in vision-language models for medical image disease classification
摘要
Medical artificial intelligence (AI) systems, particularly multimodal vision-language models (VLM), often exhibit intersectional biases, with models systematically less confident in diagnosing marginalised patient subgroups and producing higher rates of missed diagnoses. Current fairness interventions frequently fail to address these gaps or compromise overall diagnostic performance to achieve statistical parity. In this study, we developed Cross-Modal Alignment Consistency (CMAC-MMD), a training framework that standardises diagnostic certainty across intersectional patient subgroups without requiring sensitive demographic data during clinical inference. We evaluated this approach using 10,015 skin lesion images (HAM10000) with external validation on 12,000 images (BCN20000), and 10,000 fundus images for glaucoma detection (Harvard-FairVLMed), stratifying performance by intersectional age, gender, and race attributes. In the dermatology cohort, the proposed method reduced the overall intersectional missed diagnosis gap (difference in True Positive Rate, ΔTPR) from 0.50 to 0.26 while improving the Area Under the Curve (AUC) from 0.94 to 0.97 compared to standard training. Similarly, for glaucoma screening, the method reduced ΔTPR from 0.41 to 0.31, achieving a better AUC of 0.72 (vs. 0.71 baseline). This provides a methodological foundation toward clinical decision support systems that are both accurate and perform more equitably across diverse patient subgroups, without increasing privacy risks during inference.