Automated Coding Utterances Toward Chinese Course Core Competence with Large Language Models
摘要
Automated class utterance coding is a powerful tool for the evaluation of Chinese class utterances in the context of modern education, which can be seen as a multi-dimension single-label text classification task. However, due to the scarcity of data, the implicitness of Chinese language learning content, and the insufficient representation abilities of traditional machine learning methods, automated class utterance coding remains a challenging task. Therefore, we propose a framework based on a domain-specific large language model (LLM) trained on utterances from Chinese classes with well-designed prompts. The framework consists of several training strategies: data augmentation (DA) with external knowledge for background knowledge awareness, Chain-of-Thought (CoT) for reasoning ability improvement, and curriculum learning (CL) for improving coding ability from easy to hard dimensions progressively. Experiments show the effectiveness of the proposed method on the task of Chinese course core competence dimension coding for Chinese class utterances.