<p>Prior to the development of multimodal approaches, sentiment analysis of internet memes predominantly relied on single-modality techniques. With advancements in deep learning, it is now feasible to extract and integrate features from both text and image modalities, enabling more accurate sentiment prediction and classification for memes. This study presents a novel sentiment classification framework for internet memes, leveraging an enhanced ConvNeXt model combined with tensor product fusion. The ConvNeXt model is optimized through a re-parameterization structure to improve visual feature extraction from memes, while the BERT model is utilized to obtain textual features. These features are subsequently fused using a tensor fusion module to enhance sentiment prediction accuracy. Experiments are conducted on the Memotion, Memotion2, and Memotion3 datasets, with performance comparisons made against image-only, text-only, and multimodal fusion models. Experimental findings indicate that the proposed method achieves notable improvements in overall performance and classification accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion classification in internet memes utilizing enhanced ConvNeXt and tensor fusion

  • Weijun Gao,
  • Xiaoxuan Zhao

摘要

Prior to the development of multimodal approaches, sentiment analysis of internet memes predominantly relied on single-modality techniques. With advancements in deep learning, it is now feasible to extract and integrate features from both text and image modalities, enabling more accurate sentiment prediction and classification for memes. This study presents a novel sentiment classification framework for internet memes, leveraging an enhanced ConvNeXt model combined with tensor product fusion. The ConvNeXt model is optimized through a re-parameterization structure to improve visual feature extraction from memes, while the BERT model is utilized to obtain textual features. These features are subsequently fused using a tensor fusion module to enhance sentiment prediction accuracy. Experiments are conducted on the Memotion, Memotion2, and Memotion3 datasets, with performance comparisons made against image-only, text-only, and multimodal fusion models. Experimental findings indicate that the proposed method achieves notable improvements in overall performance and classification accuracy.