Contrastive Diffusion Generative Adversarial Network for Generalized Zero-Shot Learning
摘要
Generative models have been frequently used as a sample augmentation tool to tackle Generalized Zero-Shot Learning (GZSL). Generative ZSL methods aim to produce diverse and high-fidelity visual features. They do this from limited category semantics. However, on one hand, mainstream generative model families, such as GANs, VAEs, and diffusion models, cannot balance sampling speed, quality, and diversity effectively, which reduces their utility in GZSL. On the other hand, the extreme information shift between visual and semantic spaces severely restricts the discriminative and generalization abilities of the generated features. To address these challenges, we propose a novel hybrid generative model named Contrastive Diffusion Generative Adversarial Network (CD-GAN) for GZSL. As a hybrid model of the diffusion model and Generative Adversarial Network (GAN), CD-GAN incorporates the merits of the diffusion model in terms of the diversity and quality of generated features, as well as the merits of GAN in terms of feature generation efficiency. CD-GAN potentially aligns visuals and semantics by unifying real and generated features into a common embedding space and then conducting contrastive learning guided by between-class prior knowledge in this space. Moreover, we strengthen visual-semantic alignment by introducing a classification regularization. This technique further guarantees that the embedding encodes corresponding semantics. Compared to baseline models, CD-GAN achieves promising performance, as demonstrated by extensive experiments on four popular datasets.