Exploring Latent Evolving Ability in Test Equating and Its Effects on Final Rankings
摘要
This paper investigates to what extent test equalisation methods are affected by changes in the population distribution of the ability. This problem arises when the subjects can repeat the test. The scenario considered here entails two test administrations at different times and increasing ability of the subjects. Additionally, this research question becomes especially relevant when the final output of this process is a unique merit ranking. Indeed, building a merit ranking can be interpreted as a classification problem, in which the goal is to correctly classify a subject in the ranking based on their true ability. To answer these questions, we conduct a simulation study comparing concurrent calibration with the classical item response theory (IRT) linking parameter methods.