<p>Rainfall time series are essential for hydrological and climate studies; however, data scarcity remains a persistent challenge that compromises the reliability of analyses and modeling. This study compares the performance of regression based models and machine learning (ML) methods for gap filling in monthly rainfall data from Northern Minas Gerais, Brazil, a region characterized by high climatic variability and limited monitoring infrastructure. Ten missing data levels, ranging from 5% to 50%, were simulated using historical records from seven rainfall stations. Four regression based approaches, Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), and Regression Based Weighting (RWR), were evaluated alongside three ML methods, Random Forest (RF), Support Vector Machines (SVM), and k Nearest Neighbors (KNN). Model performance was assessed using Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R<sup>2</sup>), with nonparametric statistical tests applied to identify significant differences. The results showed that regression and weighting based models consistently outperformed ML methods across all missing data levels. RWR and RW achieved the best overall performance, with the lowest RMSE and SMAPE values and the highest NSE and R<sup>2</sup> values, and no statistically significant differences were observed between them. MLR performed well only at moderate levels of missing data, with significant performance degradation beyond 40% of missing values. Among ML methods, RF and SVM showed intermediate performance, while KNN and SLR yielded the weakest results. These findings highlight that weighting schemes based on the relevance and proximity of neighboring stations provide more robust rainfall estimates than complex nonlinear models in data scarce and highly variable climatic regions.</p> Graphical Abstract <p></p> <p>The graphical abstract presents the workflow used to evaluate the performance of gap-filling techniques in monthly rainfall time series from Northern Minas Gerais, Brazil. The process begins with historical rainfall data, in which artificial gaps ranging from 5% to 50% are systematically introduced to simulate missing values. These incomplete time series are then processed using seven gap filling models that include both statistical and machine learning approaches, namely Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), Regression Based Weighting (RWR), Random Forest (RF), Support Vector Machines (SVM), and K Nearest Neighbors (KNN). Each technique is applied iteratively across 100 simulations to ensure statistical reliability. The performance of the reconstructed series is evaluated using four metrics, Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R<sup>2</sup>). The flowchart visually synthesizes this stepwise approach, covering data preprocessing, gap simulation, model application, and performance comparison, and highlights the systematic strategy adopted in the study to identify the most accurate gap filling models under different levels of missing data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Filling Data Gaps: Comparing Regression-based Models and Machine Learning Methods for Rainfall Time Series Reconstruction

  • Anderson de Oliveira Pinheiro,
  • Demetrius David da Silva,
  • Michel Castro Moreira,
  • Clívia Dias Coelho,
  • Laura Thebit de Almeida,
  • Selma Alves Abrahão

摘要

Rainfall time series are essential for hydrological and climate studies; however, data scarcity remains a persistent challenge that compromises the reliability of analyses and modeling. This study compares the performance of regression based models and machine learning (ML) methods for gap filling in monthly rainfall data from Northern Minas Gerais, Brazil, a region characterized by high climatic variability and limited monitoring infrastructure. Ten missing data levels, ranging from 5% to 50%, were simulated using historical records from seven rainfall stations. Four regression based approaches, Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), and Regression Based Weighting (RWR), were evaluated alongside three ML methods, Random Forest (RF), Support Vector Machines (SVM), and k Nearest Neighbors (KNN). Model performance was assessed using Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R2), with nonparametric statistical tests applied to identify significant differences. The results showed that regression and weighting based models consistently outperformed ML methods across all missing data levels. RWR and RW achieved the best overall performance, with the lowest RMSE and SMAPE values and the highest NSE and R2 values, and no statistically significant differences were observed between them. MLR performed well only at moderate levels of missing data, with significant performance degradation beyond 40% of missing values. Among ML methods, RF and SVM showed intermediate performance, while KNN and SLR yielded the weakest results. These findings highlight that weighting schemes based on the relevance and proximity of neighboring stations provide more robust rainfall estimates than complex nonlinear models in data scarce and highly variable climatic regions.

Graphical Abstract

The graphical abstract presents the workflow used to evaluate the performance of gap-filling techniques in monthly rainfall time series from Northern Minas Gerais, Brazil. The process begins with historical rainfall data, in which artificial gaps ranging from 5% to 50% are systematically introduced to simulate missing values. These incomplete time series are then processed using seven gap filling models that include both statistical and machine learning approaches, namely Regional Weighting (RW), Simple Linear Regression (SLR), Multiple Linear Regression (MLR), Regression Based Weighting (RWR), Random Forest (RF), Support Vector Machines (SVM), and K Nearest Neighbors (KNN). Each technique is applied iteratively across 100 simulations to ensure statistical reliability. The performance of the reconstructed series is evaluated using four metrics, Root Mean Square Error (RMSE), Symmetric Mean Absolute Percentage Error (SMAPE), Nash Sutcliffe Efficiency (NSE), and the coefficient of determination (R2). The flowchart visually synthesizes this stepwise approach, covering data preprocessing, gap simulation, model application, and performance comparison, and highlights the systematic strategy adopted in the study to identify the most accurate gap filling models under different levels of missing data.