<p>Writing is a critical educational task because it encompasses so many skills necessary for the modern world, including vocabulary and grammar acquisition, critical thinking, adapting to different audiences, and determining how best to communicate one’s ideas. However, written assignments are notoriously time-consuming for teachers to grade, and timely feedback is critical for students’ learning. Automated evaluation can provide quick student feedback while easing the manual evaluation burden for teachers. Current machine learning-based methods of evaluating student textual responses have met with varying degrees of success. One main challenge in training these models is the scarcity of student-generated data. Large volumes of training data are needed to create accurate models, and few educational tasks are large enough. To overcome this data scarcity issue, text augmentation techniques have been used to balance and expand the data set so that classification models can be trained with higher accuracy, providing more useful feedback for teachers and students. This paper examines the performance of text augmentation using two Large Language Models (LLMs) to provide supplemental texts for training models for classifying student answers in English and French educational tasks. Our results show that text generation can dramatically improve model performance on small data sets over simple self-augmentation, especially when the LLM is set to generate more varied responses.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparing Text Augmentation by GPT-3.5 and Llama3 for Evaluating Student Responses

  • Keith Cochran,
  • Clayton Cohn,
  • Jean Francois Rouet,
  • Peter Hastings

摘要

Writing is a critical educational task because it encompasses so many skills necessary for the modern world, including vocabulary and grammar acquisition, critical thinking, adapting to different audiences, and determining how best to communicate one’s ideas. However, written assignments are notoriously time-consuming for teachers to grade, and timely feedback is critical for students’ learning. Automated evaluation can provide quick student feedback while easing the manual evaluation burden for teachers. Current machine learning-based methods of evaluating student textual responses have met with varying degrees of success. One main challenge in training these models is the scarcity of student-generated data. Large volumes of training data are needed to create accurate models, and few educational tasks are large enough. To overcome this data scarcity issue, text augmentation techniques have been used to balance and expand the data set so that classification models can be trained with higher accuracy, providing more useful feedback for teachers and students. This paper examines the performance of text augmentation using two Large Language Models (LLMs) to provide supplemental texts for training models for classifying student answers in English and French educational tasks. Our results show that text generation can dramatically improve model performance on small data sets over simple self-augmentation, especially when the LLM is set to generate more varied responses.