Subjective questions are crucial to assess students’ ability to analyze, synthesize, evaluate and create knowledge. In the massive online education scenarios, the manually scoring of subjective questions is time-consuming. Instead, it could be supported by the task of Short Answer Grading in Natural Language Process. However, it is worth noting that most existing automatic scoring system does not perform well on domain-specific and long questions. In this paper we address the challenges of automated short answer grading (ASAG) by proposing a novel scoring approach that strategically integrates a fine-tuned large language model (LLM), a neural network (NN) for feature extraction, and an answer-question relevance assessment module (RELEVANCE). Our method effectively scores student responses based on a set of predefined rubrics and reference answers. Our experiments on the ASAP-SAS dataset demonstrate that our method achieves an average Quadratic Weighted Kappa (QWK) score of 0.797, surpassing current state-of-the-art AutoSAS model, particularly excelling in longer tasks with a 11.9% improvement. Overall, our proposed method offers a robust solution for subjective question grading, ultimately contributing to more efficient educational assessment in a rapidly evolving learning environment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FusionASAG: An LLM-Enhanced Automatic Short Answer Grading Model for Subjective Questions in Online Education

  • He Zheng,
  • Qing Sun,
  • Qiushuo Li,
  • Yunxin Liu,
  • Yuanxin Ouyang,
  • Qinghua Cao

摘要

Subjective questions are crucial to assess students’ ability to analyze, synthesize, evaluate and create knowledge. In the massive online education scenarios, the manually scoring of subjective questions is time-consuming. Instead, it could be supported by the task of Short Answer Grading in Natural Language Process. However, it is worth noting that most existing automatic scoring system does not perform well on domain-specific and long questions. In this paper we address the challenges of automated short answer grading (ASAG) by proposing a novel scoring approach that strategically integrates a fine-tuned large language model (LLM), a neural network (NN) for feature extraction, and an answer-question relevance assessment module (RELEVANCE). Our method effectively scores student responses based on a set of predefined rubrics and reference answers. Our experiments on the ASAP-SAS dataset demonstrate that our method achieves an average Quadratic Weighted Kappa (QWK) score of 0.797, surpassing current state-of-the-art AutoSAS model, particularly excelling in longer tasks with a 11.9% improvement. Overall, our proposed method offers a robust solution for subjective question grading, ultimately contributing to more efficient educational assessment in a rapidly evolving learning environment.