<p>Multimodal learning aims to utilize the complementary information of multiple data modalities to improve the generalization performance. However, incomplete observations of multimodal data often hinder the learning of complementary information due to missing modalities. While many approaches have been proposed to address the incompleteness of test data, few methods can flexibly handle arbitrary missing modalities in both training and testing phases. In this paper, we propose a general framework called topologically consistent prototype network (TCPN) for effective and flexible solution of multimodal learning with arbitrarily missing modalities, particularly in scenarios involving two or more distinct modalities. The proposed framework establishes cross-modal associations based on the fundamental assumption of topological consistency, which posits that geometric relationships between samples and prototypes are preserved across modalities. Additionally, it introduces prototype networks to facilitate missing modal inference by employing a linear combination of modal prototypes and topological weights. For classification, multiple classifiers based on prototype learning are constructed and trained in a multi-task framework to improve feature representation and classification accuracy. Experimental results on multiple multimodal datasets demonstrate the superiority of the proposed framework over existing methods, particularly in small sample learning and imbalanced learning scenarios.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topologically Consistent Prototype Network for Incomplete Multimodal Learning

  • Yang Wang,
  • Xu-Yao Zhang,
  • Cheng-Lin Liu

摘要

Multimodal learning aims to utilize the complementary information of multiple data modalities to improve the generalization performance. However, incomplete observations of multimodal data often hinder the learning of complementary information due to missing modalities. While many approaches have been proposed to address the incompleteness of test data, few methods can flexibly handle arbitrary missing modalities in both training and testing phases. In this paper, we propose a general framework called topologically consistent prototype network (TCPN) for effective and flexible solution of multimodal learning with arbitrarily missing modalities, particularly in scenarios involving two or more distinct modalities. The proposed framework establishes cross-modal associations based on the fundamental assumption of topological consistency, which posits that geometric relationships between samples and prototypes are preserved across modalities. Additionally, it introduces prototype networks to facilitate missing modal inference by employing a linear combination of modal prototypes and topological weights. For classification, multiple classifiers based on prototype learning are constructed and trained in a multi-task framework to improve feature representation and classification accuracy. Experimental results on multiple multimodal datasets demonstrate the superiority of the proposed framework over existing methods, particularly in small sample learning and imbalanced learning scenarios.