Item position effects are typically present in international large-scale assessments (ILSAs) and can cause bias when computerized adaptive testing (CAT) is used. This study introduces an extension of the continuous calibration strategy (CCS) that aims to control for item position effects in CAT by incorporating a balanced statistical design to rotate clusters. The proposed method performed well in a simulation study with ILSA conditions (N = 2000). The CCS led to virtually unbiased ability estimates independent of the size of the item position effects across five cycles of a hypothetical ILSA test. Beginning with the second test cycle, the mean square error (MSE) of the CCS was smaller than its unbalanced version and smaller than random item selection if item position effects were present. With regard to the item response theory model used in conjunction with the CCS, the MSE of the ability estimates for the 2PL was lower from Test Cycle 3 on, but for Test Cycles 1 and 2, the MSE for the 1PL was lower. The results on the bias and MSE of the ability estimates also show that considering three clusters in the CCS is sufficient to control for item position effects in ILSA conditions. To conclude, the CCS is capable of avoiding bias in ability estimates due to item position effects, which would otherwise result when CAT is used. Additionally, due to its capability of calibrating new items during operational testing, the CCS can help to reduce the effort connected with the field trials conducted in ILSAs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Controlling for Item Position Effects When Adaptive Testing Is Used in Large-Scale Assessments

  • Andreas Frey,
  • Aron Fink

摘要

Item position effects are typically present in international large-scale assessments (ILSAs) and can cause bias when computerized adaptive testing (CAT) is used. This study introduces an extension of the continuous calibration strategy (CCS) that aims to control for item position effects in CAT by incorporating a balanced statistical design to rotate clusters. The proposed method performed well in a simulation study with ILSA conditions (N = 2000). The CCS led to virtually unbiased ability estimates independent of the size of the item position effects across five cycles of a hypothetical ILSA test. Beginning with the second test cycle, the mean square error (MSE) of the CCS was smaller than its unbalanced version and smaller than random item selection if item position effects were present. With regard to the item response theory model used in conjunction with the CCS, the MSE of the ability estimates for the 2PL was lower from Test Cycle 3 on, but for Test Cycles 1 and 2, the MSE for the 1PL was lower. The results on the bias and MSE of the ability estimates also show that considering three clusters in the CCS is sufficient to control for item position effects in ILSA conditions. To conclude, the CCS is capable of avoiding bias in ability estimates due to item position effects, which would otherwise result when CAT is used. Additionally, due to its capability of calibrating new items during operational testing, the CCS can help to reduce the effort connected with the field trials conducted in ILSAs.