<p>This study examined the air quality in Byrnihat, situated near the border of Assam and Meghalaya, a region affected by rapid industrialisation and automobile traffic. To understand the seasonal and diurnal changes in the study, are air pollutants data consist of particulate matter, gaseous pollutants and volatile organic compounds and meteorological parameters were also considered. Air Quality Index (AQI) was calculated using the guidelines of the Central Pollution Control Board (CPCB). To make the data more stable and reliable for the model prediction, the raw data was pre-processed by removing the missing and duplicate values and performing data transformation to reduce the skewness and kurtosis. Machine Learning models were used to predict and classify the AQI values. KNN was the best regression model (R<sup>2</sup> = 0.99, RMSE = 10.15). In classification, Random Forest had the best accuracy (100%), followed by LightGBM and HistGradientBoosting, which had an accuracy of approximately 98.5%. The feature importance analysis indicated that PM<sub>2.5</sub>, PM<sub>10</sub>, and NO<sub>x</sub> play a major role in the prediction of AQI. From the industry, it was evident that machine learning models can be successfully used in the air quality index prediction, in the industrial regions also, where the emissions are complex and non-linear, since the source of emission is from different sectors with wide ranges.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prediction and classification of air quality index in an industrial region using machine learning models

  • K. Jyothsna,
  • Vasavi Ravuri,
  • Sekharamahanti S Nandini,
  • P. Kanchanamala,
  • R. Udhaya,
  • P. Kalpana,
  • Christo George

摘要

This study examined the air quality in Byrnihat, situated near the border of Assam and Meghalaya, a region affected by rapid industrialisation and automobile traffic. To understand the seasonal and diurnal changes in the study, are air pollutants data consist of particulate matter, gaseous pollutants and volatile organic compounds and meteorological parameters were also considered. Air Quality Index (AQI) was calculated using the guidelines of the Central Pollution Control Board (CPCB). To make the data more stable and reliable for the model prediction, the raw data was pre-processed by removing the missing and duplicate values and performing data transformation to reduce the skewness and kurtosis. Machine Learning models were used to predict and classify the AQI values. KNN was the best regression model (R2 = 0.99, RMSE = 10.15). In classification, Random Forest had the best accuracy (100%), followed by LightGBM and HistGradientBoosting, which had an accuracy of approximately 98.5%. The feature importance analysis indicated that PM2.5, PM10, and NOx play a major role in the prediction of AQI. From the industry, it was evident that machine learning models can be successfully used in the air quality index prediction, in the industrial regions also, where the emissions are complex and non-linear, since the source of emission is from different sectors with wide ranges.