<p>Traditional deep learning models such as convolutional neural networks (CNNs), which capture localized features, and long short-term memory networks (LSTMs), which focus on long-term dependencies, often face challenges in achieving higher accuracy for time series prediction tasks. To address this limitation, this study proposes a hybrid deep learning model that integrates CNN, LSTM, the reptile search algorithm (RSA), and eXtreme Gradient Boosting (XGB) for pollutant concentration forecasting. Initially, the raw pollutant concentration data undergoes cleaning and normalization via a Min–Max scaler. The processed sequences are then separately fed into LSTM and CNN models to extract weighted features. RSA is applied to optimize these features, while XGB computes feature importance scores, quantifying the contribution of each selected feature to the predictive performance. The proposed model predicts pollutants such as <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23940_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="46" /> </InlineMediaObject> <EquationSource Format="TEX">\(PM_{2.5}\)</EquationSource> </InlineEquation>, CO, SO<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23940_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(_2\)</EquationSource> </InlineEquation>, and NO<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23940_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="8" /> </InlineMediaObject> <EquationSource Format="TEX">\(_2\)</EquationSource> </InlineEquation> up to ten days in advance for urban Indian settings. Comparative evaluations against benchmark models—including Transformer, CNN, BiLSTM, BiRNN, ANN, and BiGRU—demonstrate that the hybrid approach yields consistently superior accuracy and robustness. The hybrid model achieves substantially lower errors and higher <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41598_2025_23940_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(R^2\)</EquationSource> </InlineEquation> scores across all pollutants, validating its reliability for long-horizon air quality forecasting.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hybrid approach leveraging meta-heuristic and ensemble learning for time-sensitive prediction of pollutant concentrations

  • Priya Kansal,
  • Jatin Bedi,
  • Sushma Jain

摘要

Traditional deep learning models such as convolutional neural networks (CNNs), which capture localized features, and long short-term memory networks (LSTMs), which focus on long-term dependencies, often face challenges in achieving higher accuracy for time series prediction tasks. To address this limitation, this study proposes a hybrid deep learning model that integrates CNN, LSTM, the reptile search algorithm (RSA), and eXtreme Gradient Boosting (XGB) for pollutant concentration forecasting. Initially, the raw pollutant concentration data undergoes cleaning and normalization via a Min–Max scaler. The processed sequences are then separately fed into LSTM and CNN models to extract weighted features. RSA is applied to optimize these features, while XGB computes feature importance scores, quantifying the contribution of each selected feature to the predictive performance. The proposed model predicts pollutants such as \(PM_{2.5}\) , CO, SO \(_2\) , and NO \(_2\) up to ten days in advance for urban Indian settings. Comparative evaluations against benchmark models—including Transformer, CNN, BiLSTM, BiRNN, ANN, and BiGRU—demonstrate that the hybrid approach yields consistently superior accuracy and robustness. The hybrid model achieves substantially lower errors and higher \(R^2\) scores across all pollutants, validating its reliability for long-horizon air quality forecasting.