Multimodal learning has demonstrated promising advantages over single-modal approaches in the diagnosis of skin lesions. However, these methods often suffer from significant accuracy degradation when encountering missing modalities, hindering their clinical deployment. In this paper, we introduce a novel and effective framework, I2M2Net, for incomplete multimodal learning, focusing on adaptively and progressively mining knowledge about modal feature-aware combinations. Specifically, one branch conducts normal classification using the original complete multimodal features extracted by heterogeneous modal encoders, while another branch shares the same structures and weights, designed to perform self-distillation with masked modality combinations. These combinations are imposed on the complete features using two masking strategies simultaneously: 1) random dropout of modality (i.e. inter-modal feature masking) to simulate different missing modality combinations and foster combination-invariant dependencies, and 2) randomly mask patches of the remaining modal features (i.e. intra-modal feature masking) to promote combination-specific representations. Additionally, we design a combination-based curriculum learning (CCL) algorithm to identify weak combinations and progressively guide our network to facilitate incomplete modality learning on challenging combinations. This is achieved by adaptively adjusting the probabilities of masking based on the consistency between the complete combination and other combinations. Experimental results on the multimodal skin disease dataset Derm7pt demonstrate that our method outperforms other state-of-the-art approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

I2M2Net: Inter/Intra-modal Feature Masking Self-distillation for Incomplete Multimodal Skin Lesion Diagnosis

  • Ke Wang,
  • Linwei Qiu,
  • Yilan Zhang,
  • Fengying Xie

摘要

Multimodal learning has demonstrated promising advantages over single-modal approaches in the diagnosis of skin lesions. However, these methods often suffer from significant accuracy degradation when encountering missing modalities, hindering their clinical deployment. In this paper, we introduce a novel and effective framework, I2M2Net, for incomplete multimodal learning, focusing on adaptively and progressively mining knowledge about modal feature-aware combinations. Specifically, one branch conducts normal classification using the original complete multimodal features extracted by heterogeneous modal encoders, while another branch shares the same structures and weights, designed to perform self-distillation with masked modality combinations. These combinations are imposed on the complete features using two masking strategies simultaneously: 1) random dropout of modality (i.e. inter-modal feature masking) to simulate different missing modality combinations and foster combination-invariant dependencies, and 2) randomly mask patches of the remaining modal features (i.e. intra-modal feature masking) to promote combination-specific representations. Additionally, we design a combination-based curriculum learning (CCL) algorithm to identify weak combinations and progressively guide our network to facilitate incomplete modality learning on challenging combinations. This is achieved by adaptively adjusting the probabilities of masking based on the consistency between the complete combination and other combinations. Experimental results on the multimodal skin disease dataset Derm7pt demonstrate that our method outperforms other state-of-the-art approaches.