<p>This study investigates whether GPT-4 can reliably grade assignments in a design education context, where tasks are typically open-ended and lack a single correct answer. Such subjectivity often leads to inconsistent grading between human raters. Using an iterative research process, we developed a customized GPT-4 and tested its ability to deliver consistent assessments. The findings show that, after several rounds of refinement, the inter-rater reliability between GPT-4 and human assessors reached a level generally accepted in educational settings. This suggests that, with well-crafted prompts and customization, GPT-4 can serve as a reliable complement to human raters. We also observed moderate consistency in GPT-4’s grading over time, with intra-rater reliability scores ranging from 0.65 to 0.78. As consistency and comparability are key principles of reliable assessment, this study explores both whether a Custom GPT can meet these criteria and how iterative refinement supports this goal. The preliminary results suggest that, with appropriate prompting and iteration, GPT-4 may support more consistent grading in design education, offering a possible complement to human assessment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The application of GPT-4 in grading design university students’ assignment: an exploratory study

  • Qian Huang,
  • Thijs Willems,
  • King Wang Poon

摘要

This study investigates whether GPT-4 can reliably grade assignments in a design education context, where tasks are typically open-ended and lack a single correct answer. Such subjectivity often leads to inconsistent grading between human raters. Using an iterative research process, we developed a customized GPT-4 and tested its ability to deliver consistent assessments. The findings show that, after several rounds of refinement, the inter-rater reliability between GPT-4 and human assessors reached a level generally accepted in educational settings. This suggests that, with well-crafted prompts and customization, GPT-4 can serve as a reliable complement to human raters. We also observed moderate consistency in GPT-4’s grading over time, with intra-rater reliability scores ranging from 0.65 to 0.78. As consistency and comparability are key principles of reliable assessment, this study explores both whether a Custom GPT can meet these criteria and how iterative refinement supports this goal. The preliminary results suggest that, with appropriate prompting and iteration, GPT-4 may support more consistent grading in design education, offering a possible complement to human assessment.