The Role of Generative Artificial Intelligence in the Assessment of Open-Ended Questions
摘要
Evaluating open-ended questions presents a complex challenge, in contrast to the straightforward analysis of closed-ended or multiple-choice questions. While the latter offers clear-cut precision, open-ended questions dig deeper into student’s comprehension, allowing for a multifaceted assessment of their understanding. This analysis can be conducted at several levels. A character-level analysis distinguishes between typographical errors and substantive mistakes, providing a nuanced view of student errors. A semantic-level analysis evaluates the significance of student’s answers against the expected answer, bridging the gap between informal student language and formal correctness. Our study assesses the capability of GAI models, specifically BERT and GPT-3.5 Turbo, in the automated grading of such responses. We found that GPT-3.5 Turbo outperformed BERT, demonstrating a higher accuracy and fewer assessment errors. Our methodology included semantic-level analysis, which is crucial for interpreting the nuances of student responses. Despite this, the model’s training performance was challenged by an unbalanced dataset, particularly in lower-grade cases. As a result, the performance of GPT-3.5 Turbo appears promising for use in educational settings in the future. However, further research is required to address current dataset restrictions and explore cost-effective alternatives.