<p>Zero-shot learning (ZSL) aims to classify unlabeled instances by transferring knowledge from seen to unseen classes. The attribute space consists of terms describing various attributes of a class and is the most widely used auxiliary information in ZSL. Additionally, in word embedding, words are represented as vectors in a real number space, known as Word2Vec. Each class can be expressed using these Word2Vec. In this paper, we propose a multi-view deep generative dual fusion network for zero-shot learning, termed the Generative Dual Fusion Network (GDFN). Our model leverages both attribute features and Word2Vec to synthesize visual features for unseen classes through a generative model. Word2Vec complement the attribute features, providing a richer representation. We assume that different types of semantic information represent different views of the corresponding classes. The pseudo-visual features generated from the two types of semantic features are fused together. Adversarial training is then performed using real visual features and the fused pseudo-visual features. Additionally, visual features are matched with semantic features through regression networks, effectively enhancing the semantic space. Experimental results on four benchmark datasets demonstrate that our network benefits from the synthetic complementarity of the fused features, outperforming the latest ZSL methods and showing impressive performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-view deep generative dual fusion network for zero-shot learning

  • Xinxin Luo,
  • Wei Yin

摘要

Zero-shot learning (ZSL) aims to classify unlabeled instances by transferring knowledge from seen to unseen classes. The attribute space consists of terms describing various attributes of a class and is the most widely used auxiliary information in ZSL. Additionally, in word embedding, words are represented as vectors in a real number space, known as Word2Vec. Each class can be expressed using these Word2Vec. In this paper, we propose a multi-view deep generative dual fusion network for zero-shot learning, termed the Generative Dual Fusion Network (GDFN). Our model leverages both attribute features and Word2Vec to synthesize visual features for unseen classes through a generative model. Word2Vec complement the attribute features, providing a richer representation. We assume that different types of semantic information represent different views of the corresponding classes. The pseudo-visual features generated from the two types of semantic features are fused together. Adversarial training is then performed using real visual features and the fused pseudo-visual features. Additionally, visual features are matched with semantic features through regression networks, effectively enhancing the semantic space. Experimental results on four benchmark datasets demonstrate that our network benefits from the synthetic complementarity of the fused features, outperforming the latest ZSL methods and showing impressive performance.