<p>Equal treatment of groups and individuals is crucial for fair assessment and demand unbiased scoring decisions. We examined algorithmic fairness focusing on demographic disparities between groups of different gender and language use based on automatic scoring (<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\:n\:=\:\text{38,722}\)</EquationSource> </InlineEquation> text responses). We tested various combinations of semantic representations and classification methods on responses to reading comprehension items from the 2015 German PISA assessment. Classifications from the most accurate method, namely a Support Vector Machine trained with RoBERTa embeddings, exhibited no discernible gender differences, but a minor significant bias in the automatic scoring of students based on their language background. Specifically, students speaking mainly a foreign language at home received significantly higher automatic scores than their actual performance warranted, thereby gaining a relative advantage from the machine scoring system. Lower performing groups with more incorrect responses tend to receive more correct scores because incorrect responses are generally less likely to be recognized. Differences are particularly evident at the item level, where we identified several factors that promote algorithmic unfairness such as scoring accuracy, student performance, linguistic diversity of text responses, and the psychometrically determined item difficulty.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Algorithmic Fairness in Automatic Short Answer Scoring

  • Nico Andersen,
  • Julia Mang,
  • Frank Goldhammer,
  • Fabian Zehner

摘要

Equal treatment of groups and individuals is crucial for fair assessment and demand unbiased scoring decisions. We examined algorithmic fairness focusing on demographic disparities between groups of different gender and language use based on automatic scoring ( \(\:n\:=\:\text{38,722}\) text responses). We tested various combinations of semantic representations and classification methods on responses to reading comprehension items from the 2015 German PISA assessment. Classifications from the most accurate method, namely a Support Vector Machine trained with RoBERTa embeddings, exhibited no discernible gender differences, but a minor significant bias in the automatic scoring of students based on their language background. Specifically, students speaking mainly a foreign language at home received significantly higher automatic scores than their actual performance warranted, thereby gaining a relative advantage from the machine scoring system. Lower performing groups with more incorrect responses tend to receive more correct scores because incorrect responses are generally less likely to be recognized. Differences are particularly evident at the item level, where we identified several factors that promote algorithmic unfairness such as scoring accuracy, student performance, linguistic diversity of text responses, and the psychometrically determined item difficulty.