Recent advances in multi-label classification (MLC) have achieved promising performance based on visual representations from convolutional neural networks. Since images for MLC usually contain complicated and abundant content, utilizing only visual information is insufficient and may deviate from global context. In this paper, we integrate textual representations, denoted as semantic abstractions, with visual features to comprehensively characterize the image content. With the multi-modal information, a novel Visual-Semantic (Vise) Transformer architecture is designed for investigating the ability of the proposed abstractive semantic representations. Furthermore, we inject global context information into each local visual features, which attempts to capture the significant and concrete expressions of regions under the premise of the whole image. Empirically, extensive experiments and comparisons are conducted on the challenging benchmarks of MSCOCO and NUS-WIDE datasets. The results demonstrate the superiority of our proposed method compared with prior works with significant performance improvement for multi-label classification task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic Abstractions for Multi-label Classification

  • Xiaomei Wang,
  • Xiaohua Xuan,
  • Qing Xu,
  • Hua Cai,
  • Weilin Shen

摘要

Recent advances in multi-label classification (MLC) have achieved promising performance based on visual representations from convolutional neural networks. Since images for MLC usually contain complicated and abundant content, utilizing only visual information is insufficient and may deviate from global context. In this paper, we integrate textual representations, denoted as semantic abstractions, with visual features to comprehensively characterize the image content. With the multi-modal information, a novel Visual-Semantic (Vise) Transformer architecture is designed for investigating the ability of the proposed abstractive semantic representations. Furthermore, we inject global context information into each local visual features, which attempts to capture the significant and concrete expressions of regions under the premise of the whole image. Empirically, extensive experiments and comparisons are conducted on the challenging benchmarks of MSCOCO and NUS-WIDE datasets. The results demonstrate the superiority of our proposed method compared with prior works with significant performance improvement for multi-label classification task.