<p>As advancements in novel biomarker-based algorithms and models accelerate their use in disease risk prediction, it is crucial to evaluate these models within the context of their intended clinical application. Prediction models output the absolute risk of disease; subsequently, patient counseling and shared decision-making are based on the estimated individual risk and cost-benefit assessment. The overall impact of the application is referred to as clinical utility, which received significant attention and desire to incorporate into model assessment lately. The classic Brier score is a popular measure of prediction accuracy; however, it is insufficient for effectively assessing clinical utility. To address this limitation, we propose a class of weighted Brier scores that aligns with the decision-theoretic framework of clinical utility. Additionally, we decompose the weighted Brier score into discrimination and calibration components, and we link the weighted Brier score to the <i>H</i> measure, which has been proposed as an alternative to the area under the receiver operating characteristic curve. This theoretical link to the <i>H</i> measure further supports our weighting method and underscores the essential elements of discrimination and calibration in risk prediction evaluation. The practical use of the weighted Brier score as an overall summary is demonstrated using data from a prostate cancer study.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Weighted Brier Score—An Overall Summary Measure for Risk Prediction Models with Clinical Utility Consideration

  • Kehao Zhu,
  • Yingye Zheng,
  • Kwun Chuen Gary Chan

摘要

As advancements in novel biomarker-based algorithms and models accelerate their use in disease risk prediction, it is crucial to evaluate these models within the context of their intended clinical application. Prediction models output the absolute risk of disease; subsequently, patient counseling and shared decision-making are based on the estimated individual risk and cost-benefit assessment. The overall impact of the application is referred to as clinical utility, which received significant attention and desire to incorporate into model assessment lately. The classic Brier score is a popular measure of prediction accuracy; however, it is insufficient for effectively assessing clinical utility. To address this limitation, we propose a class of weighted Brier scores that aligns with the decision-theoretic framework of clinical utility. Additionally, we decompose the weighted Brier score into discrimination and calibration components, and we link the weighted Brier score to the H measure, which has been proposed as an alternative to the area under the receiver operating characteristic curve. This theoretical link to the H measure further supports our weighting method and underscores the essential elements of discrimination and calibration in risk prediction evaluation. The practical use of the weighted Brier score as an overall summary is demonstrated using data from a prostate cancer study.