I2M2Net: Inter/Intra-modal Feature Masking Self-distillation for Incomplete Multimodal Skin Lesion Diagnosis
摘要
Multimodal learning has demonstrated promising advantages over single-modal approaches in the diagnosis of skin lesions. However, these methods often suffer from significant accuracy degradation when encountering missing modalities, hindering their clinical deployment. In this paper, we introduce a novel and effective framework, I2M2Net, for incomplete multimodal learning, focusing on adaptively and progressively mining knowledge about modal feature-aware combinations. Specifically, one branch conducts normal classification using the original complete multimodal features extracted by heterogeneous modal encoders, while another branch shares the same structures and weights, designed to perform self-distillation with masked modality combinations. These combinations are imposed on the complete features using two masking strategies simultaneously: 1) random dropout of modality (i.e. inter-modal feature masking) to simulate different missing modality combinations and foster combination-invariant dependencies, and 2) randomly mask patches of the remaining modal features (i.e. intra-modal feature masking) to promote combination-specific representations. Additionally, we design a combination-based curriculum learning (CCL) algorithm to identify weak combinations and progressively guide our network to facilitate incomplete modality learning on challenging combinations. This is achieved by adaptively adjusting the probabilities of masking based on the consistency between the complete combination and other combinations. Experimental results on the multimodal skin disease dataset Derm7pt demonstrate that our method outperforms other state-of-the-art approaches.