Multi-view deep generative dual fusion network for zero-shot learning
摘要
Zero-shot learning (ZSL) aims to classify unlabeled instances by transferring knowledge from seen to unseen classes. The attribute space consists of terms describing various attributes of a class and is the most widely used auxiliary information in ZSL. Additionally, in word embedding, words are represented as vectors in a real number space, known as Word2Vec. Each class can be expressed using these Word2Vec. In this paper, we propose a multi-view deep generative dual fusion network for zero-shot learning, termed the Generative Dual Fusion Network (GDFN). Our model leverages both attribute features and Word2Vec to synthesize visual features for unseen classes through a generative model. Word2Vec complement the attribute features, providing a richer representation. We assume that different types of semantic information represent different views of the corresponding classes. The pseudo-visual features generated from the two types of semantic features are fused together. Adversarial training is then performed using real visual features and the fused pseudo-visual features. Additionally, visual features are matched with semantic features through regression networks, effectively enhancing the semantic space. Experimental results on four benchmark datasets demonstrate that our network benefits from the synthetic complementarity of the fused features, outperforming the latest ZSL methods and showing impressive performance.