Purpose <p>Application of computerized adaptive testing (CAT) can improve the assessment of patient-reported health outcomes by reducing patient burden. We aimed to reduce patient burden of CATs further by optimizing a standard error reduction stopping rule (SER; minimum change in SE(<i>θ</i>) after each CAT step). </p> Methods <p>We extracted PROMIS Anxiety and Depressive Symptoms CAT responses (mean age = 13.7, male = 50.3%) from the Dutch-Flemish PROMIS Assessment Center and estimated theta levels (<i>θ</i>) and standard errors (SE(<i>θ</i>)) for each step. The default stopping rules were a minimum/maximum of 4/12 items administered, respectively, or a minimum precision of SE(<i>θ</i>) &lt; 0.32. We imposed increasing SER thresholds (0.01–0.20) and compared the following outcome criteria: mean efficiency of the CAT (M<sub>efficiency</sub>; 1 − SE<i>(θ)</i><sup>2</sup>/n<sub>items</sub>), mean number of items administered (Mn<sub>items</sub>), the mean SE(<i>θ</i>) of all respondents (M<sub>SE(θ)</sub>), and mean T-score difference compared to default stopping rules (M<sub><i>∆T</i></sub>). </p> Results <p>Default stopping rules showed a mean efficiency of 0.88 and1.27, Mn<sub>items</sub> = 9.98 and8.13, and M<sub>SE(θ)</sub> = 0.36 and0.38 for respectively the Anxiety and Depressive Symptoms item banks. We optimized the SER value with a differential efficiency function, resulting in shorter, more efficient CATs (Anxiety: mean efficiency = 1.08, Mn<sub>items</sub> = 5.58, M<sub>SE(θ)</sub> = 4.24, M<sub><i>∆T</i></sub> = 0.04; Depressive Symptoms: mean efficiency = 1.45, Mn<sub>items</sub> = 4.79, M<sub>SE(θ)</sub> = 4.15, M<sub><i>∆T</i></sub> = 0.58). For participants reporting no problems, this results in fewer items administered, but a decrease in measurement accuracy and biased T-scores, which may be relevant depending on the goal of assessment. </p> Conclusions <p>We conclude that the current approach allows us to determine an optimal SER threshold that improves measurement efficiency, especially when floor/ceiling effects are present in the target population. The threshold values will vary depending on the <i>θ</i> distribution of the target population and the IRT model parameters.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reducing patient burden of PROMs in healthcare through advanced computerized adaptive testing stopping rules

  • Michiel A. J. Luijten,
  • Benjamin D. Schalet,
  • Leo D. Roorda,
  • Lotte Haverman,
  • Caroline B. Terwee

摘要

Purpose

Application of computerized adaptive testing (CAT) can improve the assessment of patient-reported health outcomes by reducing patient burden. We aimed to reduce patient burden of CATs further by optimizing a standard error reduction stopping rule (SER; minimum change in SE(θ) after each CAT step).

Methods

We extracted PROMIS Anxiety and Depressive Symptoms CAT responses (mean age = 13.7, male = 50.3%) from the Dutch-Flemish PROMIS Assessment Center and estimated theta levels (θ) and standard errors (SE(θ)) for each step. The default stopping rules were a minimum/maximum of 4/12 items administered, respectively, or a minimum precision of SE(θ) < 0.32. We imposed increasing SER thresholds (0.01–0.20) and compared the following outcome criteria: mean efficiency of the CAT (Mefficiency; 1 − SE(θ)2/nitems), mean number of items administered (Mnitems), the mean SE(θ) of all respondents (MSE(θ)), and mean T-score difference compared to default stopping rules (M∆T).

Results

Default stopping rules showed a mean efficiency of 0.88 and1.27, Mnitems = 9.98 and8.13, and MSE(θ) = 0.36 and0.38 for respectively the Anxiety and Depressive Symptoms item banks. We optimized the SER value with a differential efficiency function, resulting in shorter, more efficient CATs (Anxiety: mean efficiency = 1.08, Mnitems = 5.58, MSE(θ) = 4.24, M∆T = 0.04; Depressive Symptoms: mean efficiency = 1.45, Mnitems = 4.79, MSE(θ) = 4.15, M∆T = 0.58). For participants reporting no problems, this results in fewer items administered, but a decrease in measurement accuracy and biased T-scores, which may be relevant depending on the goal of assessment.

Conclusions

We conclude that the current approach allows us to determine an optimal SER threshold that improves measurement efficiency, especially when floor/ceiling effects are present in the target population. The threshold values will vary depending on the θ distribution of the target population and the IRT model parameters.