<p>Recent advances in Text-to-Image (T2I) diffusion models enable highly realistic image generation from text. However, long and intricate descriptions often struggle to provide precise controls. To address this, we propose <i>TexStFusion</i> (TEXtural, STructural, TEXtual feature FUSION), a method that adds conditional controls to pre-trained T2I models. Unlike existing approaches relying on visual cues, we introduce composite maps, which fuse texture and structure-text maps derived from <i>TextureNet</i> and <i>StructureNet</i> encoders. This integration occurs without fine-tuning the T2I model, preserving prior knowledge. Our method achieves <b>25%</b> better FID, <b>33%</b> better SSIM, and <b>5%</b> better CLIP-T scores with a dataset of just 30k images, in the best case.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TexStFusion : a controllable diffusion model using textural, structural, and textual feature fusion

  • Suhas Hegde,
  • Aruna Tiwari

摘要

Recent advances in Text-to-Image (T2I) diffusion models enable highly realistic image generation from text. However, long and intricate descriptions often struggle to provide precise controls. To address this, we propose TexStFusion (TEXtural, STructural, TEXtual feature FUSION), a method that adds conditional controls to pre-trained T2I models. Unlike existing approaches relying on visual cues, we introduce composite maps, which fuse texture and structure-text maps derived from TextureNet and StructureNet encoders. This integration occurs without fine-tuning the T2I model, preserving prior knowledge. Our method achieves 25% better FID, 33% better SSIM, and 5% better CLIP-T scores with a dataset of just 30k images, in the best case.