<p>Generating high-quality, semantically consistent images from text descriptions remains a challenging task in computer vision. Current methods often struggle with effectively integrating textual information into the image generation process, resulting in images that lack realism or contain significant artifacts. To address these issues, we propose SDeep, a novel framework utilizing a generative adversarial network (GAN) architecture with a channel attention mechanism. SDeep deepens the text-to-image fusion process through stacked deepening blocks (SD blocks) and enhances image detail through multilayer channel attention (MLCA). Extensive experiments on the CUB and COCO datasets demonstrate that SDeep outperforms state-of-the-art methods in terms of image quality and semantic alignment with text descriptions. Our approach not only generates more realistic images but also better preserves the semantic consistency between text and generated images, marking a significant advancement in text-to-image synthesis. Code can be found at <a href="https://github.com/zxcnmmmmm/SDeep">https://github.com/zxcnmmmmm/SDeep</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stacked deep fusion GAN for enhanced text-to-image generation

  • Wenli Chen,
  • Yaqi Sun,
  • Paul L. Rosin,
  • YuKun Lai

摘要

Generating high-quality, semantically consistent images from text descriptions remains a challenging task in computer vision. Current methods often struggle with effectively integrating textual information into the image generation process, resulting in images that lack realism or contain significant artifacts. To address these issues, we propose SDeep, a novel framework utilizing a generative adversarial network (GAN) architecture with a channel attention mechanism. SDeep deepens the text-to-image fusion process through stacked deepening blocks (SD blocks) and enhances image detail through multilayer channel attention (MLCA). Extensive experiments on the CUB and COCO datasets demonstrate that SDeep outperforms state-of-the-art methods in terms of image quality and semantic alignment with text descriptions. Our approach not only generates more realistic images but also better preserves the semantic consistency between text and generated images, marking a significant advancement in text-to-image synthesis. Code can be found at https://github.com/zxcnmmmmm/SDeep.