A-calibration: assessment of prediction models for survival data under censoring
摘要
Evaluating the performance of predictive models for survival is essential before they can be trusted for real-world applications and decision making. While good measures such as the C-index are available for model discrimination, the toolbox for model calibration is much more limited in the time-to-event setting.
The method of D-calibration was therefore an important contribution that yields a single numeric value for calibration across the available follow-up time. D-calibration consists of performing a Pearson’s goodness-of-fit test on transformed survival times. Censored survival times are handled using an imputation approach which however tends to yield a conservative test and loss of power.
MethodsIn this paper, we introduce A-calibration based on Akritas’s goodness-of-fit test which is designed specifically for censored time-to-event data. Through theoretical arguments, simulations, and a case study, we compare A- and D-calibration as measures of calibration. In the simulation study, the power of each test to reject a false null-hypothesis was assessed for varying censoring mechanisms (memoryless, uniform and zero censoring), censoring rates, and parameter values of the predictive model considered.
ResultsThe simulation study demonstrated that A-calibration had similar or superior power to D-calibration in all considered cases, and that D-calibration, unlike A-calibration, was particularly sensitive to censoring.
ConclusionsAdvantages of A-calibration compared to D-calibration have been demonstrated through theoretical considerations, a simulation study, and a case study, while no disadvantages relative to D-calibration were identified.