What Inconsistent Assessment Outcomes Reveal About Learning Heterogeneity Across Structurally Equivalent Writing Tasks
摘要
Higher education assessment practices often treat parallel writing assignments graded with identical criteria as interchangeable evidence of student learning. This study examines whether structurally equivalent writing tasks in fact function as interchangeable indicators of performance and what cross-task variation reveals about learning heterogeneity and feedback uptake. The analysis draws on two structurally equivalent writing assessments administered in two parallel sections of an upper-division undergraduate international relations elective course (N = 47). To account for differences in grading stringency across tasks, performance was evaluated using within-assessment standardized z-scores. The analysis combines correlations, internal consistency measures, Bland–Altman agreement analysis, transitions across relative performance categories, and exploratory probit models assessing whether academic background characteristics predict improvement among initially low-performing students. Despite aligned rubric criteria and comparable cognitive demands, performance across the two assessments showed only modest associations, low internal consistency, wide limits of agreement, and substantial reclassification of students’ relative standing. Standardized performance changes revealed divergent learning trajectories, with some students improving markedly and others declining. Students with higher cumulative GPAs and more completed credits were directionally more likely to move from below-mean to above-mean performance, although estimates were imprecise due to sample size. Overall, the findings indicate that structurally equivalent writing assessments do not operate as interchangeable measures of learning. Instead, cross-task variability reflects meaningful differences in how students engage with successive writing opportunities, suggesting that repeated assessments can illuminate learning heterogeneity that single-task evaluations obscure and informing course-level assessment practices aimed at supporting feedback uptake and learning transfer.