Personalized products based on individual preferences have been considered to improve personal well-being and consumer satisfaction. This approach helps reduce waste and conserve resources. With artificial intelligence enabling personalization, consumers can easily access products that match their preferences without the need for specialized knowledge or professional expertise. Advances in artificial intelligence, text-to-image models in particular, have enabled the generation of impressive images from textual descriptions. However, existing models lack the ability to generate images based on visual impressions. In this paper, we propose a text-to-image diffusion model that incorporates visual impressions into the image generation process. Our model extends the stable diffusion architecture by introducing a multi-modal input system that processes text descriptions, pattern images, and quantified visual impressions. Experimental validation confirmed the positive correlation between generated and original images across multiple impression metrics, demonstrating the model’s effectiveness in preserving impression-based characteristics. These results suggest that our approach successfully bridges the gap between textual descriptions and visual impressions in image generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generation of Clothing Patterns Based on Impressions Using Stable Diffusion

  • J. N. Htoi Sann Ja,
  • Kaede Shiohara,
  • Toshihiko Yamasaki,
  • Miyuki Toga,
  • Kensuke Tobitani,
  • Noriko Nagata

摘要

Personalized products based on individual preferences have been considered to improve personal well-being and consumer satisfaction. This approach helps reduce waste and conserve resources. With artificial intelligence enabling personalization, consumers can easily access products that match their preferences without the need for specialized knowledge or professional expertise. Advances in artificial intelligence, text-to-image models in particular, have enabled the generation of impressive images from textual descriptions. However, existing models lack the ability to generate images based on visual impressions. In this paper, we propose a text-to-image diffusion model that incorporates visual impressions into the image generation process. Our model extends the stable diffusion architecture by introducing a multi-modal input system that processes text descriptions, pattern images, and quantified visual impressions. Experimental validation confirmed the positive correlation between generated and original images across multiple impression metrics, demonstrating the model’s effectiveness in preserving impression-based characteristics. These results suggest that our approach successfully bridges the gap between textual descriptions and visual impressions in image generation.