Prediction of spontaneous preterm birth in pregnant women using machine learning
摘要
Spontaneous preterm birth (sPTB) is a significant global health concern, contributing to adverse outcomes for both pregnant women and newborns. Early identification of women with risk of sPTB is essential for mitigating these negative effects and improving maternal and neonatal health outcomes. The aim of this study is to explore the feasibility of using machine learning to predict sPTB risk and to analyze the contribution of variables.
MethodsAll data were collected retrospectively. Prediction models were developed using eight different machine learning algorithms combined with six variable selection methods. The models’ predictive performance was evaluated using area under the receiver operating characteristic curve (AUROC), area under the precision recall curve (AUPRC), accuracy, sensitivity, F1-score, positive predictive value, and negative predictive value.
ResultsA total of 1122 pregnant women, of whom 187 had preterm birth and 935 had term birth, were enrolled. The model by combining the categorical boosting algorithm and backward elimination had the best predictive performance with the highest AUROC (0.8762) and AUPRC (0.7061), and the Brier score was 0.12 on the test set. The top 5 variables for predicting sPTB risk in this study were free triiodothyronine, albumin/globulin, thyroglobulin antibody, total thyroxine, red cell volume distribution width.
ConclusionsThe machine learning model may help identify pregnant women at high risk of sPTB, and individual risk factor analysis could provide reference for clinical decision. However, as some key variables are not part of routine laboratory tests during pregnancy worldwide, the model’s generalizability and clinical applicability require further study.