Automatic short answer grading can significantly enhance the speed and fairness of grading, making it particularly valuable in areas with a shortage of teachers, such as Africa [1]. However, for most African languages it is very challenging to build automatic short answer grading systems due to the limited availability of natural language processing corpora. Furthermore, only experts can deal with the complex algorithms, required for training and fine-tuning traditional automatic short answer grading systems. Given that state-of-the-art large language models have the potential to address these problems through their growing capabilities and ease of use through prompting, particularly in zero-shot and few-shot learning, we investigated their performance for grading student answers in the African language Twi. To address the absence of a Twi corpus, we translated and validated the University of North Texas benchmark corpus [2], creating the first Twi automatic short answer grading corpus. On this corpus, we evaluated the performances of the large language models GPT-4o [3], Claude 3 Sonnet [4], and LLaMA 3 [5] as well as for comparison two more traditional approaches: a fine-tuned AfroLM and a cross-lingual M-BERT approach. Among individual models, the cross-lingual M-BERT had the best performance with a mean absolute error of 0.79 points out of 5 points, followed by fine-tuned AfroLM at 0.73 points and Claude 3 Sonnet at 1.00 points. However, combining AfroLM and M-BERT outputs achieved the lowest mean absolute error of 0.64 points, which is less than the human grader variance of 0.75 points in the original corpus [6]. Combining the outputs of the large language models GPT-4o, Claude 3 Sonnet, and LLaMA 3, obtained through few-shot learning, yielded a mean absolute error of 1.10 points.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI in Education: An Analysis of Large Language Models for Twi Automatic Short Answer Grading

  • Alex Agyemang,
  • Tim Schlippe

摘要

Automatic short answer grading can significantly enhance the speed and fairness of grading, making it particularly valuable in areas with a shortage of teachers, such as Africa [1]. However, for most African languages it is very challenging to build automatic short answer grading systems due to the limited availability of natural language processing corpora. Furthermore, only experts can deal with the complex algorithms, required for training and fine-tuning traditional automatic short answer grading systems. Given that state-of-the-art large language models have the potential to address these problems through their growing capabilities and ease of use through prompting, particularly in zero-shot and few-shot learning, we investigated their performance for grading student answers in the African language Twi. To address the absence of a Twi corpus, we translated and validated the University of North Texas benchmark corpus [2], creating the first Twi automatic short answer grading corpus. On this corpus, we evaluated the performances of the large language models GPT-4o [3], Claude 3 Sonnet [4], and LLaMA 3 [5] as well as for comparison two more traditional approaches: a fine-tuned AfroLM and a cross-lingual M-BERT approach. Among individual models, the cross-lingual M-BERT had the best performance with a mean absolute error of 0.79 points out of 5 points, followed by fine-tuned AfroLM at 0.73 points and Claude 3 Sonnet at 1.00 points. However, combining AfroLM and M-BERT outputs achieved the lowest mean absolute error of 0.64 points, which is less than the human grader variance of 0.75 points in the original corpus [6]. Combining the outputs of the large language models GPT-4o, Claude 3 Sonnet, and LLaMA 3, obtained through few-shot learning, yielded a mean absolute error of 1.10 points.