Data-driven water quality prediction using hybrid machine learning approaches for sustainable development goal 6
摘要
Sustainability Development Goal (6) emphasizes the need of providing quality water and sanitation, in order to enhance health outcomes on a global scale. Although there were several conventional procedures for detecting the quality of water but they led to high costs and time consumption. The aim of the paper was the development of an automated system to predict type of water based on different designed target classes. These classes include drinking water, outdoor bathing water and irrigation water whose values were scaled on the basis of pH, conductivity, and so on using standard water quality criteria. Raw data from various Indian regions over the past 5–6 years was collected from the Indian Government's official website with the attributes like fecal coliform, dissolved oxygen, pH value, nitrate, and conductivity. Further, it had been preprocessed, scaled, as well as augmented using KNN, feature scaling and SMOTE technique respectively to normalize the values of the water attributes and improve the classification performance. Ten machine learning algorithms and a hybrid technique were trained on the water type prediction dataset. During testing, the results showed that for the drinking water class, LGBM achieved the highest precision, and F1 score values, for the outdoor bathing class, Random Forest, XgBoost, and LGBM attained the highest F1 score while as for the irrigation water class, only XgBoost and LGBM computed the highest F1 score values of 100%. These results indicate that the automated system has the potential to classify the type of water which provides a cost-efficient and effective solution to advance the goals of Sustainable Development Goal 6.