<p>Water contamination is a worldwide issue affecting semiarid regions, such as Morocco. Increased anthropogenic and natural water pollution has led to interest in the development of novel instruments for analyzing water quality. Computation of the water quality index is usually laborious and prone to errors. Moreover, conducting experiments to ascertain the sensitivity of water quality factors is costly. Therefore, a study using five machine learning algorithms, K-nearest neighbor (KNN), artificial neural network (ANN), naive Bayes (NB), and random forest (RF), was conducted to analyze 30-year samples (1988–2017). Total phosphorus (TP), fecal coliform (FC), ammonium (NH<sub>4</sub>*), dissolved oxygen (DO), biochemical oxygen demand (BOD<sub>5</sub>), and chemical oxygen demand (COD) were used as explanatory variables. The dichotomous water quality index represents the dependent variable. The model testing and training used 80% and 20% of the data, respectively. The results enabled us to determine the most accurate and efficient model for predicting the two types of water quality target classes identified in the study. A confusion matrix and a series of statistical measurements were used to evaluate the overall performance of the generated prediction models on both training and test datasets. The random forest classifier performed better than the other models, according to the model validation findings, which included an accuracy of positive predicted value (100%), negative predicted value (98.83%), F-measure (99.33%), and Kappa index (0.987). To forecast water quality with minimal inputs, a scenario with three independent variables (PT, COD, and DBO<sub>5</sub>) identified as having a major impact on water quality was explored and evaluated. The high accuracy (above 84% of correct classification) of the four models (RF, ANN, NB, and SVM) demonstrates that it is possible to use fewer water quality parameters to forecast the water quality index. Decision makers can monitor river water more effectively using this straightforward, quick, affordable, and time-efficient model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Application of machine learning techniques to assess surface water quality in the Sebou Basin, Morocco

  • Khalid Chadli

摘要

Water contamination is a worldwide issue affecting semiarid regions, such as Morocco. Increased anthropogenic and natural water pollution has led to interest in the development of novel instruments for analyzing water quality. Computation of the water quality index is usually laborious and prone to errors. Moreover, conducting experiments to ascertain the sensitivity of water quality factors is costly. Therefore, a study using five machine learning algorithms, K-nearest neighbor (KNN), artificial neural network (ANN), naive Bayes (NB), and random forest (RF), was conducted to analyze 30-year samples (1988–2017). Total phosphorus (TP), fecal coliform (FC), ammonium (NH4*), dissolved oxygen (DO), biochemical oxygen demand (BOD5), and chemical oxygen demand (COD) were used as explanatory variables. The dichotomous water quality index represents the dependent variable. The model testing and training used 80% and 20% of the data, respectively. The results enabled us to determine the most accurate and efficient model for predicting the two types of water quality target classes identified in the study. A confusion matrix and a series of statistical measurements were used to evaluate the overall performance of the generated prediction models on both training and test datasets. The random forest classifier performed better than the other models, according to the model validation findings, which included an accuracy of positive predicted value (100%), negative predicted value (98.83%), F-measure (99.33%), and Kappa index (0.987). To forecast water quality with minimal inputs, a scenario with three independent variables (PT, COD, and DBO5) identified as having a major impact on water quality was explored and evaluated. The high accuracy (above 84% of correct classification) of the four models (RF, ANN, NB, and SVM) demonstrates that it is possible to use fewer water quality parameters to forecast the water quality index. Decision makers can monitor river water more effectively using this straightforward, quick, affordable, and time-efficient model.