<p>This study presents a comprehensive assessment of air quality prediction in Hyderabad, India, using advanced machine learning models. Daily pollutant data from 2015 to 2020, including PM<sub>2.5</sub>, PM<sub>10</sub>, NO, NO₂, NOx, NH₃, CO, SO₂, O₃, benzene, toluene, and xylene, were analyzed to develop predictive frameworks. Five ensemble models (LightGBM, RandomForest, CatBoost, Adaboost, and XGBoost) were evaluated using metrics such as MAE, RMSE, and R². RandomForest and XGBoost emerged as the top-performing models, with RandomForest achieving an R² of 0.223 and XGBoost an R² of 0.188 on the validation set. Feature importance analysis consistently identified PM<sub>10</sub>, NO₂, and O₃ as dominant predictors across models. The study’s novelty lies in its integrated approach, combining statistical distribution analysis, temporal pollutant trends, and ensemble interpretability, to enhance urban air quality forecasting. These findings demonstrate the potential of machine learning to support data-driven pollution mitigation strategies in rapidly urbanizing cities like Hyderabad.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Harnessing Machine Learning for Air Quality Prediction: A Case Study of Hyderabad

  • Guhan Velusamy,
  • Dharma Raju Akasapu,
  • Naga Ratna Kopparthi,
  • Kavitha Chandu

摘要

This study presents a comprehensive assessment of air quality prediction in Hyderabad, India, using advanced machine learning models. Daily pollutant data from 2015 to 2020, including PM2.5, PM10, NO, NO₂, NOx, NH₃, CO, SO₂, O₃, benzene, toluene, and xylene, were analyzed to develop predictive frameworks. Five ensemble models (LightGBM, RandomForest, CatBoost, Adaboost, and XGBoost) were evaluated using metrics such as MAE, RMSE, and R². RandomForest and XGBoost emerged as the top-performing models, with RandomForest achieving an R² of 0.223 and XGBoost an R² of 0.188 on the validation set. Feature importance analysis consistently identified PM10, NO₂, and O₃ as dominant predictors across models. The study’s novelty lies in its integrated approach, combining statistical distribution analysis, temporal pollutant trends, and ensemble interpretability, to enhance urban air quality forecasting. These findings demonstrate the potential of machine learning to support data-driven pollution mitigation strategies in rapidly urbanizing cities like Hyderabad.