<p>The task of text-to-image generation has matured significantly with the advancement of diffusion models; however, achieving precise control over the details and style of generated images remains a challenge. Existing methods often rely on complex text prompts to describe details while attempting to integrate style information. Nevertheless, the single-stage attention mechanism in diffusion models struggles to effectively capture multi-scale features and the relationship between style and content, resulting in feature amalgamation that compromises the quality of generated images. To address this issue, we propose a Style-Content Progressive Aggregation (SCPA) network, which integrates and aggregates multi-scale features from style images and text prompts through the coordinated design of two complementary modules. Specifically, the Style-Content Decoupling (SCD) module disentangles the style and content features of the style image, and reconstructs a learnable content template based on the extracted style features, thereby preventing the original content features from interfering with text understanding. The Style-Content Coupling (SCC) module then extracts multi-scale pixel-level content features from the text prompt and progressively integrates style elements into the template, enabling fine-grained fusion of content and style. This progressive aggregation strategy effectively enhances the quality of prior guidance provided to the diffusion model. Extensive experimental results demonstrate that the SCPA network can generate more artistically appealing images and offers a new direction for the integration of text-to-image generation models with traditional style transfer techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Style-Content progressive aggregation network with stable diffusion

  • Tiebiao Yuan,
  • Yangyang Yu,
  • Ning Ji

摘要

The task of text-to-image generation has matured significantly with the advancement of diffusion models; however, achieving precise control over the details and style of generated images remains a challenge. Existing methods often rely on complex text prompts to describe details while attempting to integrate style information. Nevertheless, the single-stage attention mechanism in diffusion models struggles to effectively capture multi-scale features and the relationship between style and content, resulting in feature amalgamation that compromises the quality of generated images. To address this issue, we propose a Style-Content Progressive Aggregation (SCPA) network, which integrates and aggregates multi-scale features from style images and text prompts through the coordinated design of two complementary modules. Specifically, the Style-Content Decoupling (SCD) module disentangles the style and content features of the style image, and reconstructs a learnable content template based on the extracted style features, thereby preventing the original content features from interfering with text understanding. The Style-Content Coupling (SCC) module then extracts multi-scale pixel-level content features from the text prompt and progressively integrates style elements into the template, enabling fine-grained fusion of content and style. This progressive aggregation strategy effectively enhances the quality of prior guidance provided to the diffusion model. Extensive experimental results demonstrate that the SCPA network can generate more artistically appealing images and offers a new direction for the integration of text-to-image generation models with traditional style transfer techniques.