Sticker images have become widely popular as potent tools for emotional expression in contemporary online chats. Existing efforts in sticker emotion recognition typically integrate both global and local information, with a particular emphasis on key regions of the images. However, sticker images often contain complex elements in addition to highlighting the main subject, which can lead to potential omission and confusion in selecting local information. To alleviate this issue, we propose ELEMO, which utilizes local connectivity and fine-grained feature to emphasize the important information of the sticker. Specifically, We leverage overlapping patch embedding to preserve the related information between patches of the sticker image. Then we combine the object segmentation algorithm with text-image similarity to enhance the reliability of local information representation. Experimental results on the SER30K dataset demonstrate the effectiveness of our method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ELEMO: Elements Focused Emotion Recognition for Sticker Images

  • Min Luo,
  • Boda Lin,
  • Binghao Tang,
  • Haolong Yan,
  • Si Li

摘要

Sticker images have become widely popular as potent tools for emotional expression in contemporary online chats. Existing efforts in sticker emotion recognition typically integrate both global and local information, with a particular emphasis on key regions of the images. However, sticker images often contain complex elements in addition to highlighting the main subject, which can lead to potential omission and confusion in selecting local information. To alleviate this issue, we propose ELEMO, which utilizes local connectivity and fine-grained feature to emphasize the important information of the sticker. Specifically, We leverage overlapping patch embedding to preserve the related information between patches of the sticker image. Then we combine the object segmentation algorithm with text-image similarity to enhance the reliability of local information representation. Experimental results on the SER30K dataset demonstrate the effectiveness of our method.