This paper addresses the increasing demand for accurate, fair, and efficient grading of essay-style assessments in higher education by integrating institutional requirements with recent advances in large language models (LLMs). We propose a pipeline emphasizing privacy, explainability, consistency, and fairness. To ensure privacy, the system operates on local servers and employs rigorous anonymization of student data. Grading events integrate task prompts, instructor guidelines, student submissions, grader commentary, and final scores into structured records, enhancing evaluation accuracy and transparency. We detail the development and validation process, fine-tuning two local LLMs using historical course data. Results demonstrate the models’ ability to effectively replicate original grading decisions while safeguarding student data. Additionally, we discuss how this framework aligns with the interests of students, educators, and policymakers. Our approach establishes a foundational methodology for responsibly integrating AI-driven grading into higher education. This contribution fosters trust among stakeholders and sets a clear direction for future implementation and research efforts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enabling Responsible LLM-Based Grading in Higher Education – Design Guidelines and a Reproducible Data Preparation Pipeline

  • Arnold F. Arz von Straussenburg,
  • Anna Wolters,
  • Timon T. Aldenhoff,
  • Dennis M. Riehle

摘要

This paper addresses the increasing demand for accurate, fair, and efficient grading of essay-style assessments in higher education by integrating institutional requirements with recent advances in large language models (LLMs). We propose a pipeline emphasizing privacy, explainability, consistency, and fairness. To ensure privacy, the system operates on local servers and employs rigorous anonymization of student data. Grading events integrate task prompts, instructor guidelines, student submissions, grader commentary, and final scores into structured records, enhancing evaluation accuracy and transparency. We detail the development and validation process, fine-tuning two local LLMs using historical course data. Results demonstrate the models’ ability to effectively replicate original grading decisions while safeguarding student data. Additionally, we discuss how this framework aligns with the interests of students, educators, and policymakers. Our approach establishes a foundational methodology for responsibly integrating AI-driven grading into higher education. This contribution fosters trust among stakeholders and sets a clear direction for future implementation and research efforts.