Vision Language Models (VLMs) have revolutionized core computer vision tasks such as classification and segmentation. However, their potential in creative fields, such as art and design, which involve numerous vision-related tasks, remains underexplored, partly due to a gap between computer vision scientists and art/design professionals. This paper aims to bridge this gap by providing a ‘window’ that empowers artists and designers to more easily utilize VLM and domain-specific datasets, thus improving both divergent and convergent thinking processes. We propose a comprehensive framework for integrating VLMs into creative workflows, focusing on style transfer, visual concept mapping, automated rendering, and curatorial storytelling. By mapping these tasks to divergent and convergent thinking modes, we demonstrate how VLMs can augment human-centered creativity. We also discuss the growing availability of domain-specific datasets–such as WikiArt and The Metropolitan Museum of Art’s Open Access Artworks [1, 2]–which enable VLMs to be fine-tuned for specialized applications in art and design. Finally, we examine ethical considerations, including bias, cultural sensitivity, and misinterpretation, to foster inclusive creative practices. This framework underscores a forward-looking approach to harnessing VLMs as complements, rather than replacements, for human ingenuity and outlines pathways for future research and applications in human-computer interaction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Aug-Creativity: Framework for Human-Centered Creativity with Vision Language Models

  • Dan Li,
  • Lei Xia,
  • Ling Fan

摘要

Vision Language Models (VLMs) have revolutionized core computer vision tasks such as classification and segmentation. However, their potential in creative fields, such as art and design, which involve numerous vision-related tasks, remains underexplored, partly due to a gap between computer vision scientists and art/design professionals. This paper aims to bridge this gap by providing a ‘window’ that empowers artists and designers to more easily utilize VLM and domain-specific datasets, thus improving both divergent and convergent thinking processes. We propose a comprehensive framework for integrating VLMs into creative workflows, focusing on style transfer, visual concept mapping, automated rendering, and curatorial storytelling. By mapping these tasks to divergent and convergent thinking modes, we demonstrate how VLMs can augment human-centered creativity. We also discuss the growing availability of domain-specific datasets–such as WikiArt and The Metropolitan Museum of Art’s Open Access Artworks [1, 2]–which enable VLMs to be fine-tuned for specialized applications in art and design. Finally, we examine ethical considerations, including bias, cultural sensitivity, and misinterpretation, to foster inclusive creative practices. This framework underscores a forward-looking approach to harnessing VLMs as complements, rather than replacements, for human ingenuity and outlines pathways for future research and applications in human-computer interaction.