SoulSearch: applying heuristic optimization to enhance text-to-image generation with personalized human-LMM collaboration
摘要
Large multimodal models (LMMs) have emerged as a transformative force in the field of artificial intelligence, offering various potential applications and opportunities. Despite significant advancements in LMM techniques, the practical application of LMMs, particularly in text-to-image generation, faces limitations due to the lack of effective collaborative designs between these models and human users. To address the limitation, we embark on an empirical study of users’ iterative usage of text-to-image composition, uncovering the challenges that users face and their corresponding expectations. Then, we introduce SoulSearch, a system designed to enhance the efficiency and personalized satisfaction of text-to-image composition. The core concept behind SoulSearch draws inspiration from heuristic optimization, where exploration and exploitation processes are integrated into each turn to guide users toward optimal editing directions. Our evaluations demonstrate SoulSearch’s advantages across various text-to-image tasks, with notable improvements in usability and usefulness metrics.