Anemia is a serious health problem characterized by low blood red blood cell counts, affecting 37% of pregnant women worldwide. The objective of this study was to predict anemia in pregnant women using machine learning and deep learning techniques; the research can contribute to prenatal care and reduce anemia in pregnant women in Peru. The methodology was based on four (4) phases: Obtaining the dataset; Preprocessing; Model training: Machine Learning (SVM, Random Forest and Gradient Boosting); and Deep Learning (CNN and MLP) and Evaluation (Accuracy, Precision, Recall and F1-Score). The feature selection process using LightGBM indicated that variables such as altitude and the poverty index of the district of residence were closely related to the diagnosis of anemia, whereas variables such as the type of pregnancy were of little significance. The model trained with CNN obtained an accuracy of 82%, recall of 92%, precision of 75%, and an F1-score of 83%, suggesting the need to use datasets with variables that consider economic and geographic aspects.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Anemia Detection Model for Pregnant Women in Peru Using Machine Learning and Deep Learning

  • Sebastian Gamarra,
  • Robert Velasquez,
  • Wilfredo Ticona

摘要

Anemia is a serious health problem characterized by low blood red blood cell counts, affecting 37% of pregnant women worldwide. The objective of this study was to predict anemia in pregnant women using machine learning and deep learning techniques; the research can contribute to prenatal care and reduce anemia in pregnant women in Peru. The methodology was based on four (4) phases: Obtaining the dataset; Preprocessing; Model training: Machine Learning (SVM, Random Forest and Gradient Boosting); and Deep Learning (CNN and MLP) and Evaluation (Accuracy, Precision, Recall and F1-Score). The feature selection process using LightGBM indicated that variables such as altitude and the poverty index of the district of residence were closely related to the diagnosis of anemia, whereas variables such as the type of pregnancy were of little significance. The model trained with CNN obtained an accuracy of 82%, recall of 92%, precision of 75%, and an F1-score of 83%, suggesting the need to use datasets with variables that consider economic and geographic aspects.