<p>In modern artificial intelligence research, multimodal learning has emerged as a crucial research area. It effectively integrates data from different modalities, allowing for a more comprehensive understanding of complex informational contexts. With the widespread application of large models, it is essential to conduct in-depth research on how to effectively fuse different data from different modalities using transformer architectures. To this end, this study proposes a hierarchical interaction multimodal model based on RoBERTa-Keyword-ViT. First, we proposed a method to effectively extract keywords, ensuring that the keyword information is not lost while preserving the semantic information of the text. Second, the modal primacy of pairs in cross-attention was discussed for the first time in this paper, and a residual primary–secondary cross-attention interaction mechanism was proposed. The proposed mechanism ensures that the features of the primary modal data are effectively preserved during the interactions among different modalities. Finally, we discussed the role of query, key, and value in the primary–secondary modality interaction and verified the advantages of using the primary modality as a carrier and the secondary modality as a query. These findings enhance the understanding of multimodal interaction mechanisms and provide an empirical basis for future research. This indicates that selecting appropriate primary–secondary modalities and their roles in attention mechanisms is crucial when designing multimodal interaction systems. Lastly, we conducted experiments on both relatively large- and small-scale public datasets. Compared with existing studies, the proposed approach achieved superior performance on slightly larger dataset leveraging the RoBERTa-Keyword-ViT model architecture.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hierarchical interaction multimodal model for feature fusion based on RoBERTa-Keyword-ViT

  • Yuanhang Wang,
  • Yonghua Zhou,
  • Min Zhong,
  • Yiduo Mei,
  • Hamido Fujita,
  • Hanan Aljuaid

摘要

In modern artificial intelligence research, multimodal learning has emerged as a crucial research area. It effectively integrates data from different modalities, allowing for a more comprehensive understanding of complex informational contexts. With the widespread application of large models, it is essential to conduct in-depth research on how to effectively fuse different data from different modalities using transformer architectures. To this end, this study proposes a hierarchical interaction multimodal model based on RoBERTa-Keyword-ViT. First, we proposed a method to effectively extract keywords, ensuring that the keyword information is not lost while preserving the semantic information of the text. Second, the modal primacy of pairs in cross-attention was discussed for the first time in this paper, and a residual primary–secondary cross-attention interaction mechanism was proposed. The proposed mechanism ensures that the features of the primary modal data are effectively preserved during the interactions among different modalities. Finally, we discussed the role of query, key, and value in the primary–secondary modality interaction and verified the advantages of using the primary modality as a carrier and the secondary modality as a query. These findings enhance the understanding of multimodal interaction mechanisms and provide an empirical basis for future research. This indicates that selecting appropriate primary–secondary modalities and their roles in attention mechanisms is crucial when designing multimodal interaction systems. Lastly, we conducted experiments on both relatively large- and small-scale public datasets. Compared with existing studies, the proposed approach achieved superior performance on slightly larger dataset leveraging the RoBERTa-Keyword-ViT model architecture.