Comparison and evaluation of machine learning models for predicting indoor PM2.5 concentrations on a large spatiotemporal scale
摘要
The underlying uncertainty associated with long-term exposure to indoor pollutants at the population level has prevented point prediction models for indoor PM2.5 from providing adequate information for large-scale applications. Moreover, physics-based prediction models are constrained by the untraceable input complexity. In this study, we predicted the large-scale spatiotemporal distributions of residential PM2.5 concentration using three data-driven models: Gaussian Process Regression (GPR), Quantile Random Forest (QRF), and Bayesian Neural Network (BNN). These three models were selected based on their established representative status within the spectrum of machine learning, ranging from “shallow” to “deep” methodologies. Our findings underscore the superior performance of the BNN model, which achieved an R2 ranging from 0.48 to 0.70 and 95% prediction interval coverage between 85% and 88% across multiple datasets. The comprehensive framework presented herein for model comparison, validation, and attribution can assist future studies in elucidating the complex nonlinear relationships between urban characteristics and indoor air pollutants, thereby providing valuable insights into urban planning, design, and policy development from the perspective of indoor PM2.5 pollution mitigation.