Text to Image Generation Based on Adaptive Attention
摘要
In our paper, we propose the Adaptive Attention-based Generative Adversarial Network (AAGAN) for text to image generation, and the modal combines the multi-layer GANs and Adaptive Attention Mechanisms to control the fine-grained image generation process at different levels. The core components of AAGAN include the Adaptive Attention Module (AAM) and Spatial-Channel Instance Normalization (SCIN). AAM can dynamically improve the attention weights depending on instructive attention standard based on the paired text-image features in both spatial and channel dimensions. In addition, we propose the loss function for spatial and channel respectively to constrain the proximity of the correlation to the instructive standard. SCIN ensures that the training process is not influenced by samples within the same batch. According to our experimental results, AAGAN can achieve high-quality image generation based on natural language descriptions.