Improving Open-Response Assessment with LearnLM
摘要
Automated grading of learners’ open responses remains challenging due to the complexity of language and the subjective nature of human evaluation. Recent advances in generative AI, particularly large language models (LLMs), offer new possibilities to improve assessment. Off-the-shelf LLMs, such as GPT-4, have been applied to this task, as well as dedicated education-oriented models, such as LearnLM. However, little is known about their effectiveness compared to general-purpose models. In this study, we evaluate GPT-4o, GPT-4-turbo, and Gemini-Pro and compare their performance to LearnLM to determine their effectiveness in assessing learning, specifically the professional development of adult tutors. We find that LearnLM outperforms other models on tasks requiring tutor learners to predict the most appropriate response to students. We hypothesize that this is due to the model’s fine-tuning on tutor-student interaction data and suggest that LearnLM may be particularly useful in scenario-based tutor training. To further improve automated assessment methods, we challenge the concept of human “ground truth” by proposing alternative validation methods. Specifically, we introduce a predictive validity method by relating open-response scores with corresponding multiple-choice scores that demonstrate statistically significant and moderate correlations, particularly with LearnLM. Our novel method demonstrates predictive validity but should be combined with additional measures to ensure a more comprehensive assessment. This study contributes an open source dataset, human annotation rubrics, and LLM prompts, to improve future assessment applications of LLMs.