Generative text-to-image platforms have become essential tools across fields such as art, design, and education. While they unlock creative potential by bridging language and vision, there is a lack of sufficient research focusing on user-related challenges with AI text-to-image platforms. This gap underscores the need for a comparative study of AI text-to-image platforms to evaluate their usability and effectiveness from the user’s perspective. This study conducts a comprehensive comparison of three mainstream platforms—DALL-E 3, MidJourney, and Stable Diffusion—through a within-subjects design involving 20 participants from diverse academic backgrounds. Participants completed two tasks per platform: replicating a reference image and creating images based on creative themes. The think-aloud method, image quality ratings (5-point Likert scale), the System Usability Scale (SUS), and semi-structured interviews were used to evaluate usability and user strategies. Results indicate MidJourney excels in image quality and also demonstrates good usability. DALL-E 3 offers slightly higher usability but with moderate image quality. In contrast, Stable Diffusion, while provides advanced customization, is hindered by a steep learning curve and technical complexity, resulting in poorest image quality. Common challenges among three platforms include inconsistent prompt execution, difficulties in describing ideas effectively, and AI artifacts. These findings offer actionable insights for improving platform usability and supporting the broader adoption of text-to-image tools.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identifying Usability Challenges in Text-to-Image AI: A Comprehensive Comparison Among Mainstream Platforms

  • Yuqing Cai,
  • Yihan Zhou,
  • Xiaoqun Yu,
  • Woojoo Kim

摘要

Generative text-to-image platforms have become essential tools across fields such as art, design, and education. While they unlock creative potential by bridging language and vision, there is a lack of sufficient research focusing on user-related challenges with AI text-to-image platforms. This gap underscores the need for a comparative study of AI text-to-image platforms to evaluate their usability and effectiveness from the user’s perspective. This study conducts a comprehensive comparison of three mainstream platforms—DALL-E 3, MidJourney, and Stable Diffusion—through a within-subjects design involving 20 participants from diverse academic backgrounds. Participants completed two tasks per platform: replicating a reference image and creating images based on creative themes. The think-aloud method, image quality ratings (5-point Likert scale), the System Usability Scale (SUS), and semi-structured interviews were used to evaluate usability and user strategies. Results indicate MidJourney excels in image quality and also demonstrates good usability. DALL-E 3 offers slightly higher usability but with moderate image quality. In contrast, Stable Diffusion, while provides advanced customization, is hindered by a steep learning curve and technical complexity, resulting in poorest image quality. Common challenges among three platforms include inconsistent prompt execution, difficulties in describing ideas effectively, and AI artifacts. These findings offer actionable insights for improving platform usability and supporting the broader adoption of text-to-image tools.