Avalon: A Human-in-the-Loop LLM Grading System with Instructor Calibration and Student Self-assessment
摘要
Open-ended assessments promote deeper learning but pose significant challenges for timely, high-quality feedback–often burdening instructors with lengthy grading processes and resulting in underutilized student feedback. We introduce Avalon, a human-in-the-loop AI grading system that integrates (1) an iterative calibration phase to align AI graders with instructor expectations and (2) a novel student self-grading and discrepancy reporting mechanism. Through rubric calibration, instructors provide corrective feedback on AI-graded samples, ensuring consistent application of grading criteria. After the AI grades all submissions, students assess their own work using the same rubric, then compare their scores with the AI’s. They submit short “discrepancy reports” for any mismatches, distinguishing between accepted differences and genuine disputes. In a pilot with 102 undergraduates, Avalon reduced instructor grading time by focusing manual review on a small subset - fewer than 16% of submissions - that were disputed by students. Moreover, students show high engagement with the feedback process, as self-grading compelled them to revisit rubric criteria and reflect on their submissions more deeply. Furthermore, the system uncovered misconceptions that might otherwise have gone undetected–prompting targeted instructor intervention. Although additional validation and larger-scale studies are needed, current preliminary findings suggest Avalon’s hybrid approach can reduce grading workloads, improve feedback effectiveness, and enhance student engagement and learning outcomes. The Avalon platform can be accessed at https://avalonlearn.com