Evaluating differential item functioning through effect size measures in logistic regression: implications for language assessment
摘要
This study investigates the consistency of statistical tests and the effectiveness of effect size measures in evaluating items with Differential Item Functioning (DIF) in a large-scale assessment. We applied a stepwise binary Logistic Regression (LR) analysis to a large-scale exam sat by 999 PhD applicants in Iran. First, we employed Likelihood Ratio Test (LRT) and Wald separately on the same dichotomous data to detect their equivalence in identifying DIF. The LRT compares two models to see whether including group membership (e.g., gender) significantly improves the model's fit to the data, while the Wald test checks whether the difference in performance between groups is statistically significant based on model coefficients. Then to evaluate the effectiveness of effect size measures, we combined Nagelkerke ΔR2 with statistical tests and, in a separate analysis, blended Δ log of odds ratio (Δ LR) effect size criterion with confidence intervals. ΔR2 estimates how much of the variation in responses is explained by group membership (similar to a percentage of explained variance), and DLR quantifies how much more or less likely one group is to answer an item correctly compared to another group. The results indicated that blending the effect size measures with statistical tests or confidence intervals was more effective in identifying practically significant DIF-flagged items than using the statistical tests alone. Furthermore, it was indicated that Δ LR was less conservative than Δ R2 in identifying practically significant DIF-flagged items. We suggest that assessment developers apply effect size measures combined with statistical tests to identify practically significant DIF-flagged items, hence ensuring assessment fairness.