<p>This study introduces an innovative approach to enhancing e-commerce product listings through subject-driven text-to-image generation, leveraging advanced AI technologies. Focused on transforming consumer first impressions, it blends personalized visual styles with online retail needs, striking a balance between standardization and customization. The research develops a unique method for image synthesis, improving upon existing AI models such as DreamBooth and Textual Inversion. This work not only equips online sellers with dynamic visual tools but also significantly enriches AI applications in e-commerce, offering both practical and academic contributions to the field. Our proposed model is evaluated based on various numerical and human-based evaluation metrics. The experimental results show that our model achieves a significant performance compared to other baseline models. Our model is further analyzed and discussed under correlation analysis, visual quality assessment, and ablation study to ensure its practical applicability and user satisfaction.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel approach with vision-language models for custom e-commerce product listings

  • Y Huynh Ngoc Nhu,
  • Quoc-Dung Nguyen,
  • Cherdsak Kingkan

摘要

This study introduces an innovative approach to enhancing e-commerce product listings through subject-driven text-to-image generation, leveraging advanced AI technologies. Focused on transforming consumer first impressions, it blends personalized visual styles with online retail needs, striking a balance between standardization and customization. The research develops a unique method for image synthesis, improving upon existing AI models such as DreamBooth and Textual Inversion. This work not only equips online sellers with dynamic visual tools but also significantly enriches AI applications in e-commerce, offering both practical and academic contributions to the field. Our proposed model is evaluated based on various numerical and human-based evaluation metrics. The experimental results show that our model achieves a significant performance compared to other baseline models. Our model is further analyzed and discussed under correlation analysis, visual quality assessment, and ablation study to ensure its practical applicability and user satisfaction.