<p>In positive and unlabeled (PU) learning problems, only positive examples are labeled. Unlabeled data contain both positive and negative examples. Studies show that positive examples of (secondary) diagnoses, and clinical conditions, such as sepsis, are present in unlabeled hospital administrative data, potentially distorting hospital reimbursement systems, and negatively affecting hospitals’ revenue and profitability. We investigate whether PU learning is suitable for improving the quality of hospital administrative data. We train three models on 313,434 hospital cases using hospital cost features: two based on the two-step “spy” approach and one using a robust PU learning method. For model evaluation, we rely exclusively on positive examples due to the PU setting. To further assess model performance, we perform an external validity check: We relabel unlabeled sepsis cases, derive new sepsis rates, and compare them to those reported in medical record review studies. All models identify true positives well in unseen data. External validity checks show, however, that only the robust PU learner effectively discriminates between positives and negatives in the unlabeled data, yielding new sepsis rates within the range of sepsis rates reported in medical record review studies. PU learning can improve the quality of hospital administrative data, but its effectiveness depends strongly on the choice of learning approach and classifier. The output of a PU learner can potentially improve hospital reimbursement systems, hospital revenue and profitability management, and sensitivity analyses in healthcare management science, health economics, health services research, and disease surveillance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Positive and unlabeled learning from hospital administrative data: a novel approach to identify sepsis cases

  • Justus Vogel,
  • Johannes Cordier

摘要

In positive and unlabeled (PU) learning problems, only positive examples are labeled. Unlabeled data contain both positive and negative examples. Studies show that positive examples of (secondary) diagnoses, and clinical conditions, such as sepsis, are present in unlabeled hospital administrative data, potentially distorting hospital reimbursement systems, and negatively affecting hospitals’ revenue and profitability. We investigate whether PU learning is suitable for improving the quality of hospital administrative data. We train three models on 313,434 hospital cases using hospital cost features: two based on the two-step “spy” approach and one using a robust PU learning method. For model evaluation, we rely exclusively on positive examples due to the PU setting. To further assess model performance, we perform an external validity check: We relabel unlabeled sepsis cases, derive new sepsis rates, and compare them to those reported in medical record review studies. All models identify true positives well in unseen data. External validity checks show, however, that only the robust PU learner effectively discriminates between positives and negatives in the unlabeled data, yielding new sepsis rates within the range of sepsis rates reported in medical record review studies. PU learning can improve the quality of hospital administrative data, but its effectiveness depends strongly on the choice of learning approach and classifier. The output of a PU learner can potentially improve hospital reimbursement systems, hospital revenue and profitability management, and sensitivity analyses in healthcare management science, health economics, health services research, and disease surveillance.