By rationally embedding word information, we present a novel single-stage generative adversarial network architecture (Dual Attention Single-Stage GAN, DAS-GAN) to produce high-quality textual description images. The framework incorporates two innovative attention modules: a spatial-aware-attention module at the word level and a channel-aware-attention module at the character level. Through the process of modelling the significance of individual words within a specific sentence, these two modules collaborate to allocate greater weights to word vectors directly linked to image characteristics. Furthermore, we have developed a novel feature fusion module that comprises two independent word-level affine modules. Compared to the previous affine transform layer that only uses sentence information, the word-level affine transform layer in the new fusion module not only fuses sentence information during the information fusion process, but also enables the image sub-regions to be effectively fused with the most relevant word information. Extensive trials have been carried out on the difficult MS COCO dataset and the CUB bird dataset. The experimental findings demonstrate that our suggested DAS-GAN markedly enhances both the FID and IS metrics, yielding images with superior visual fidelity compared to conventional approaches. Furthermore, our approach demonstrates superior performance compared to earlier multi-stage and single-stage approaches that just rely on sentence information.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DAS-GAN: A Dual Attention Single-Stage GAN for Text-to-Image Generation

  • Lingling Zi,
  • Xiaolin Chen,
  • Xin Cong

摘要

By rationally embedding word information, we present a novel single-stage generative adversarial network architecture (Dual Attention Single-Stage GAN, DAS-GAN) to produce high-quality textual description images. The framework incorporates two innovative attention modules: a spatial-aware-attention module at the word level and a channel-aware-attention module at the character level. Through the process of modelling the significance of individual words within a specific sentence, these two modules collaborate to allocate greater weights to word vectors directly linked to image characteristics. Furthermore, we have developed a novel feature fusion module that comprises two independent word-level affine modules. Compared to the previous affine transform layer that only uses sentence information, the word-level affine transform layer in the new fusion module not only fuses sentence information during the information fusion process, but also enables the image sub-regions to be effectively fused with the most relevant word information. Extensive trials have been carried out on the difficult MS COCO dataset and the CUB bird dataset. The experimental findings demonstrate that our suggested DAS-GAN markedly enhances both the FID and IS metrics, yielding images with superior visual fidelity compared to conventional approaches. Furthermore, our approach demonstrates superior performance compared to earlier multi-stage and single-stage approaches that just rely on sentence information.