<p>Machine learning classifiers trained on imbalanced healthcare datasets often exhibit bias, leading to poor performance on critical cases. The cardiotocography (CTG) dataset exemplifies this issue, where misclassification of pathological cases arises due to both class imbalance and non-optimal probability thresholds. Statistical analysis suggests refining classification thresholds, but this approach has been largely overlooked in CTG data research. To address these challenges, we propose a multifusion method integrating undersampling, threshold-moving optimization, and ensemble classifiers to enhance classification precision while maintaining computational efficiency. Applied to a CTG dataset of 502 cases from Czech Technical University and University Hospital Brno, our method showed significant improvements in identifying pathological cases. While baseline models correctly classified only about 2 out of 11 cases per test, our approach achieved 76.92, 75, and 41.67% precision, accurately identifying 9, 9, and 3 cases out of 12, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing cardiotocography classification via ensemble learning and threshold optimization

  • Lingping Kong,
  • Václav Snášel,
  • Zhonghai Bai,
  • Dominik Vilimek,
  • Seyedali Mirjalili,
  • Jeng-Shyang Pan,
  • Jitka Horakova,
  • Radek Martinek,
  • Radana Vilimkova Kahankova

摘要

Machine learning classifiers trained on imbalanced healthcare datasets often exhibit bias, leading to poor performance on critical cases. The cardiotocography (CTG) dataset exemplifies this issue, where misclassification of pathological cases arises due to both class imbalance and non-optimal probability thresholds. Statistical analysis suggests refining classification thresholds, but this approach has been largely overlooked in CTG data research. To address these challenges, we propose a multifusion method integrating undersampling, threshold-moving optimization, and ensemble classifiers to enhance classification precision while maintaining computational efficiency. Applied to a CTG dataset of 502 cases from Czech Technical University and University Hospital Brno, our method showed significant improvements in identifying pathological cases. While baseline models correctly classified only about 2 out of 11 cases per test, our approach achieved 76.92, 75, and 41.67% precision, accurately identifying 9, 9, and 3 cases out of 12, respectively.