In this paper, we propose a novel Conditioning Augmentation and Attention Mechanism Generative Adversarial Network (CAAM-GAN) for text-to-image generation. The Conditioning Augmentation(CA) module enhances the text semantic features by increasing the feature data corresponding to text descriptions to improve the diversity of generated images. We added the Efficient Channel Attention(ECA) module to the discriminator to learn the correlation between image features through cross-channel interaction to improve the quality of generated images. In addition, to improve semantic consistency between text descriptions and generated images, we incorporate the Convolutional Block Attention Module(CBAM) into the generator. By using the CBAM to adjust the output features from both channel and spatial dimensions, the generator focuses on the important features of the text descriptions, suppresses the unnecessary features, and generates higher-quality images with text-image semantic consistency and diversity. Our method is validated on the CaltechUCSD Birds 200 (CUB) dataset and the Microsoft Common Objects in Context (COCO) dataset. The experimental results demonstrate the effectiveness and superiority of our method. In terms of both subjective and objective evaluation, the results of our method surpass the existing state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CAAM-GAN: Text to Image Generation with Conditioning Augmentation and Attention Mechanism

  • Hongjie Cai,
  • Jun Shang,
  • Wenxin Yu,
  • Lu Che,
  • Zhiqiang Zhang,
  • Ning Jiang,
  • Kang Xu,
  • Jun Gong,
  • Peng Chen

摘要

In this paper, we propose a novel Conditioning Augmentation and Attention Mechanism Generative Adversarial Network (CAAM-GAN) for text-to-image generation. The Conditioning Augmentation(CA) module enhances the text semantic features by increasing the feature data corresponding to text descriptions to improve the diversity of generated images. We added the Efficient Channel Attention(ECA) module to the discriminator to learn the correlation between image features through cross-channel interaction to improve the quality of generated images. In addition, to improve semantic consistency between text descriptions and generated images, we incorporate the Convolutional Block Attention Module(CBAM) into the generator. By using the CBAM to adjust the output features from both channel and spatial dimensions, the generator focuses on the important features of the text descriptions, suppresses the unnecessary features, and generates higher-quality images with text-image semantic consistency and diversity. Our method is validated on the CaltechUCSD Birds 200 (CUB) dataset and the Microsoft Common Objects in Context (COCO) dataset. The experimental results demonstrate the effectiveness and superiority of our method. In terms of both subjective and objective evaluation, the results of our method surpass the existing state-of-the-art methods.