User interface (UI) design often requires many rounds of feedback from heuristic evaluations by human experts, which is resource consuming and can be inconsistent. Many contributions have been made towards using large language models (LLMs) to automate UI design; however, existing works only diagnose a singular UI screen at a time. For heuristic evaluation, it is crucial to analyze multiple UI screens simultaneously to consider internal consistency across an entire application. To fill this gap, we propose an AI-powered method, AIHeurEval, built on the GPT-4o mini MLLMs to analyze multiple UI screens simultaneously and perform automated heuristic evaluation, focusing on the heuristic of internal consistency, including color, layout, typography, and tone. The experiment results show that AIHeurEval with One-Shot Chain-of-Thought (CoT) prompting has a significant improvement on overall performance. It achieved a 43.25% increase in identified issues, 51.35% increase in proposed recommendations, and 14.28% decrease in hallucinations compared to GPT-4o mini against the human baseline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AIHeurEval: Generating Heuristic Evaluations on Multiple UI Screens with Multimodal Large Language Models

  • Yuekai Wang,
  • Franceska Xhakaj

摘要

User interface (UI) design often requires many rounds of feedback from heuristic evaluations by human experts, which is resource consuming and can be inconsistent. Many contributions have been made towards using large language models (LLMs) to automate UI design; however, existing works only diagnose a singular UI screen at a time. For heuristic evaluation, it is crucial to analyze multiple UI screens simultaneously to consider internal consistency across an entire application. To fill this gap, we propose an AI-powered method, AIHeurEval, built on the GPT-4o mini MLLMs to analyze multiple UI screens simultaneously and perform automated heuristic evaluation, focusing on the heuristic of internal consistency, including color, layout, typography, and tone. The experiment results show that AIHeurEval with One-Shot Chain-of-Thought (CoT) prompting has a significant improvement on overall performance. It achieved a 43.25% increase in identified issues, 51.35% increase in proposed recommendations, and 14.28% decrease in hallucinations compared to GPT-4o mini against the human baseline.