This work presents a Text-to-Image model-based product image background generation method, Concept-Edge Fusion, which generates a high-quality background for a product freely through text description while maintaining the details of the product and without deforming the edges. Existing methods for generating product backgrounds often lead to semantic misunderstanding and edge expansion, which greatly undermines the quality of the generated image. Semantic misunderstanding represents the apparent difference between the main subject of the generated image and that of the reference product image when the model generates a new image based on the text description. Edge expansion refers to the apparent changes in the contour/shape of the input product. To solve these problems, we introduce Concept-Inject and Edge-Control modules to help the Text-to-Image model better generate the background guided by the text description. The concept-inject module prevents the model from semantic misunderstanding about the given product, and the edge-control module ensures that the product edges are not expanded when completing the product background. Extensive experiments demonstrate that our method can better perform background generation for products without changing the semantics, shape, and details of the original products. We also construct a dataset to evaluate the text-guided product background generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Concept-Edge Fusion: Background Generation for Product Presentation Based on Text-to-Image Model

  • Pengfei Deng,
  • Tianjiao Zhang,
  • Weize Quan,
  • Hanyu Wang,
  • Qinglin Lu,
  • Zhifeng Li,
  • Dong-Ming Yan

摘要

This work presents a Text-to-Image model-based product image background generation method, Concept-Edge Fusion, which generates a high-quality background for a product freely through text description while maintaining the details of the product and without deforming the edges. Existing methods for generating product backgrounds often lead to semantic misunderstanding and edge expansion, which greatly undermines the quality of the generated image. Semantic misunderstanding represents the apparent difference between the main subject of the generated image and that of the reference product image when the model generates a new image based on the text description. Edge expansion refers to the apparent changes in the contour/shape of the input product. To solve these problems, we introduce Concept-Inject and Edge-Control modules to help the Text-to-Image model better generate the background guided by the text description. The concept-inject module prevents the model from semantic misunderstanding about the given product, and the edge-control module ensures that the product edges are not expanded when completing the product background. Extensive experiments demonstrate that our method can better perform background generation for products without changing the semantics, shape, and details of the original products. We also construct a dataset to evaluate the text-guided product background generation.