Exploring the Role of Large Language Models in Evaluating Argumentative Writing in Military School Education
摘要
The Professiographic profile of a military school student is built through observations made during the student’s training. One of these observations is the evaluation of argumentative writing. This study investigates the use of Large Language Models (LLM), such as GPT-4, in evaluating this argumentative writing, focusing on opinion articles. The aim is to explore the potential of LLM to provide instant and personalized feedback at different stages of writing, as well as to assess their effectiveness compared to human evaluators. Using a detailed rubric that covers everything from topic choice to bibliographical references, initial results indicate that GPT-4 can consistently evaluate technical and structural aspects of writing, offering reliable feedback, especially in the References category. However, its conservative approach may underestimate the quality of the articles, suggesting the need for human oversight. The study also highlights the challenges faced by GPT-4 with more subtle and contextual elements of opinion writing, evidenced by variability in precision and low recall in recognizing complete works. These findings underscore the evolving role of LLM as supplementary tools in education, which should be integrated with human judgment to enhance argumentative writing and critical thinking in academic settings.