A pilot study on forecasting PM2.5 in oil field industrial locality with statistical and AI/ML approaches
摘要
Air pollution is a leading cause of mortality in the developing world, with industrial regions experiencing particularly high levels of fine particulate matter (PM2.5). The Assam Oil Field Region, a significant industrial hub, is characterized by elevated PM2.5 concentrations, yet predictive modeling studies in such environments remain scarce. This study presents a preliminary exploration of statistical and machine learning techniques, aiming to identify potential tools for future environmental monitoring, in oil field localities. Using a limited dataset (n = 109), we modelled PM2.5 levels based on multiple predictor variables, including meteorological parameters (temperature, humidity, wind speed, and wind direction). We employed multiple linear regression alongside six machine learning models—Linear Regression, Decision Tree Regression, Support Vector Regression, KNN Regression, XGBoost Regression (XGBR), and Random Forest Regression, and one Feed Forward Deep Learning model. Additionally, Principal Component Analysis was employed to investigate the impact of dimensionality reduction on model performance. Among the models, XGBR consistently demonstrated superior predictive performance, achieving high R2 scores (0.78 in non-PCA, 0.59 in PCA) and minimal error metrics (MAE: 6.75, MSE: 62.72, RMSE: 7.92). Its robustness across scenarios and ability to handle complex relationships efficiently make it particularly effective compared to other ensemble methods like Random Forest Regression. While initial outcomes suggest the potential of machine learning models, especially XGBoost, for PM2.5 predictions in industrial contexts, further validation using larger and more diverse datasets is necessary. These early findings may serve as a basis for future work on data-driven air quality forecasting in oil field regions.
Graphical Abstract