While usability aspects play a crucial role in socio-technical systems, proper usability evaluations are often neglected during development due to time and financial constraints or the lack of availability of usability experts. To assess the suitability of generative AI as a usability evaluator, we conducted a heuristic evaluation comparing GPT-4 to a group of seven human evaluators with backgrounds in human-computer interaction, computer science or behavioral science. As performance measures we used the number of identified usability issues and severity ratings as well as the nature and quality of the problem description. No significant differences were found regarding identification rates of usability issues or severity ratings except for the overall identification rate. Thematic analysis, however, showed that the problem descriptions differed qualitatively in specificity, coverage, clarity and insightfulness. GPT-4 generated more detailed, very clear and insightful problem descriptions of a broader coverage than human evaluators. Our results provide initial evidence for the suitability of GPT-4 as a usability expert. However, the results should be viewed with caution as sample size was small and included only one expert. Future studies are needed, including more experienced usability experts and more GPT-4 cases to properly test for quantitative differences.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Heuristic Evaluation of Kadi4Mat Through Human Evaluators and GPT-4

  • Annika Meinecke,
  • David Heidrich,
  • Katharina Dworatzyk,
  • Sabine Theis

摘要

While usability aspects play a crucial role in socio-technical systems, proper usability evaluations are often neglected during development due to time and financial constraints or the lack of availability of usability experts. To assess the suitability of generative AI as a usability evaluator, we conducted a heuristic evaluation comparing GPT-4 to a group of seven human evaluators with backgrounds in human-computer interaction, computer science or behavioral science. As performance measures we used the number of identified usability issues and severity ratings as well as the nature and quality of the problem description. No significant differences were found regarding identification rates of usability issues or severity ratings except for the overall identification rate. Thematic analysis, however, showed that the problem descriptions differed qualitatively in specificity, coverage, clarity and insightfulness. GPT-4 generated more detailed, very clear and insightful problem descriptions of a broader coverage than human evaluators. Our results provide initial evidence for the suitability of GPT-4 as a usability expert. However, the results should be viewed with caution as sample size was small and included only one expert. Future studies are needed, including more experienced usability experts and more GPT-4 cases to properly test for quantitative differences.