Background <p>Infertility, defined as the inability to conceive after 12 months of regular unprotected sexual intercourse, is an important public health concern with significant social, psychological, and marital consequences. In this study, infertility was based on self-reported time to conception rather than clinical diagnosis; no clinical examination or male factor assessment was conducted, and outcome misclassification is therefore possible. Despite its importance, community-based evidence in Ethiopia remains limited. This study aimed to estimate the prevalence of infertility, identify factors associated with infertility, and evaluate the predictive performance of selected machine learning models among married women in North Shewa Zone, Oromia, Ethiopia.</p> Methods <p>A community-based cross-sectional study was conducted among 829 randomly selected married women from three districts. Infertility status was self-reported based on time to conception. Multivariable logistic regression with Firth penalization was used to identify factors associated with infertility, and adjusted odds ratios (AORs) with 95% confidence intervals (CIs) were reported. For predictive analysis, machine learning models—including Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM)—were developed using an 80/20 train–test split with 5-fold cross-validation applied to the training data. Class imbalance was addressed using resampling techniques. Model performance was evaluated using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC).</p> Results <p>The prevalence of infertility was 12.7%. Factors significantly associated with infertility included district of residence, age at marriage, irregular menstrual cycles, history of abortion, family history of infertility, number of sexual partners, cigarette smoking, and women’s occupation. In predictive modelling, Random Forest achieved the highest AUC (0.89), followed by XGBoost (0.88) and SVM (0.81), compared with logistic regression (AUC = 0.69). However, differences in predictive performance were modest. Given the relatively small number of infertility events, there is potential for model instability and overfitting. Variable importance analysis from the Random Forest model identified women’s occupation, education level, family history of infertility, and menstrual cycle pattern as important predictors.</p> Conclusion <p>Infertility among married women in North Shewa Zone is associated with multiple socio-demographic, reproductive, and behavioral factors. However, findings should be interpreted cautiously due to the cross-sectional design, self-reported outcome, absence of clinical verification, and limited number of events. Although machine learning models demonstrated relatively higher predictive performance, results are based on internal validation only and may not generalize to other populations. Logistic regression remains appropriate for inference, while machine learning approaches may be useful for exploratory prediction. Further studies with larger sample sizes, improved outcome measurement, and external validation are required before practical application.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prevalence, associated factors, and machine learning–based prediction of infertility among married women in North Shewa Zone, Ethiopia: a cross-sectional study

  • Abate Tadesse Zeleke,
  • Tadesse Ayele Belachew,
  • Merga Abdissa Aga,
  • Degemu Sahlu Asebe,
  • Rebik Shukure Beyane

摘要

Background

Infertility, defined as the inability to conceive after 12 months of regular unprotected sexual intercourse, is an important public health concern with significant social, psychological, and marital consequences. In this study, infertility was based on self-reported time to conception rather than clinical diagnosis; no clinical examination or male factor assessment was conducted, and outcome misclassification is therefore possible. Despite its importance, community-based evidence in Ethiopia remains limited. This study aimed to estimate the prevalence of infertility, identify factors associated with infertility, and evaluate the predictive performance of selected machine learning models among married women in North Shewa Zone, Oromia, Ethiopia.

Methods

A community-based cross-sectional study was conducted among 829 randomly selected married women from three districts. Infertility status was self-reported based on time to conception. Multivariable logistic regression with Firth penalization was used to identify factors associated with infertility, and adjusted odds ratios (AORs) with 95% confidence intervals (CIs) were reported. For predictive analysis, machine learning models—including Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM)—were developed using an 80/20 train–test split with 5-fold cross-validation applied to the training data. Class imbalance was addressed using resampling techniques. Model performance was evaluated using accuracy, sensitivity, specificity, and area under the receiver operating characteristic curve (AUC).

Results

The prevalence of infertility was 12.7%. Factors significantly associated with infertility included district of residence, age at marriage, irregular menstrual cycles, history of abortion, family history of infertility, number of sexual partners, cigarette smoking, and women’s occupation. In predictive modelling, Random Forest achieved the highest AUC (0.89), followed by XGBoost (0.88) and SVM (0.81), compared with logistic regression (AUC = 0.69). However, differences in predictive performance were modest. Given the relatively small number of infertility events, there is potential for model instability and overfitting. Variable importance analysis from the Random Forest model identified women’s occupation, education level, family history of infertility, and menstrual cycle pattern as important predictors.

Conclusion

Infertility among married women in North Shewa Zone is associated with multiple socio-demographic, reproductive, and behavioral factors. However, findings should be interpreted cautiously due to the cross-sectional design, self-reported outcome, absence of clinical verification, and limited number of events. Although machine learning models demonstrated relatively higher predictive performance, results are based on internal validation only and may not generalize to other populations. Logistic regression remains appropriate for inference, while machine learning approaches may be useful for exploratory prediction. Further studies with larger sample sizes, improved outcome measurement, and external validation are required before practical application.