Vision and language pre-training (VLP) has demonstrated significant effectiveness across various vision-language (V+L) tasks. Traditional unsupervised VLP methods predominantly utilize a unimodal masking model, typically comprising a visual encoder and a textual en- coder. These models often restrict masking to textual data only. Consequently, such a unimodal approach limits the effective integration, or ’fusibility,’ of features extracted from each modality. Addressing this limitation, this paper introduces the Masked Multimodal Autoencoder (MME). MME innovatively enhances feature fusibility by optimizing mutual information across masked, cross-modal data. This is achieved through the integration of a dynamic masking strategy, applicable to both visual and textual encoders, coupled with a novel pre-training task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Only Alignment and Reconstruction: Masked Multimodal Autoencoders for Vison-Language Pre-training

  • Chong Jiang,
  • Jinglei Tang

摘要

Vision and language pre-training (VLP) has demonstrated significant effectiveness across various vision-language (V+L) tasks. Traditional unsupervised VLP methods predominantly utilize a unimodal masking model, typically comprising a visual encoder and a textual en- coder. These models often restrict masking to textual data only. Consequently, such a unimodal approach limits the effective integration, or ’fusibility,’ of features extracted from each modality. Addressing this limitation, this paper introduces the Masked Multimodal Autoencoder (MME). MME innovatively enhances feature fusibility by optimizing mutual information across masked, cross-modal data. This is achieved through the integration of a dynamic masking strategy, applicable to both visual and textual encoders, coupled with a novel pre-training task.