A benchmark dataset with human validation for AI-assisted technical answer evaluation
摘要
Grading open-ended technical responses has been a longstanding issue in higher education. Despite advances in automated assessment, existing approaches often rely on holistic scoring, weakly validated annotations, and cognitive alignment, limiting pedagogical reliability and classroom adoption. To facilitate reliable and rubric-based automated evaluation in the Data Structures and Algorithms course, this study presents DSA-RubricEval, a pedagogically grounded and preliminarily reliability-assessed dataset. The dataset constitutes the primary contribution of this work, providing a structured benchmark aligned with Bloom’s taxonomy and validated through multi-rater Interclass Correlation Coefficient ICC analysis. The dataset consists of twelve expert-designed, Bloom-aligned, questions scored on multiple rubric dimensions using an ordinal scale. Five independent evaluators scored student responses, enabling rigorous validation of human judgement based on the Intraclass Correlation Coefficient (ICC) analysis. The results indicate good to excellent average-measure reliability across most rubric dimensions, justifying the use of aggregated human scores as aggregated reference labels for automated assessment. Automated scoring was explored as a proof of concept to demonstrate the applicability of the dataset for AI-assisted assessment and formulated as an ordinal, rubric-level prediction task and evaluated using pedagogically motivated agreement measures, showing high tolerance-based agreement with human judges. The proposed research identifies sources of assessor subjectivity and explores methods to mitigate them, while reducing the workload of grading as well as correlate with learning outcomes. Moreover, the proposed technique supports lower-order cognitive skills but is not best suited for higher-order cognitive tasks. Overall, this study introduces a preliminarily reliability-assessed, rubric-based benchmark dataset intended to support exploratory research on pedagogically meaningful AI-assisted assessment of open-ended technical responses.