<p>Taiwan has introduced specific policies, such as the 2050 net-zero goal and the 2024 carbon fee, to support global climate action and encourage carbon reduction efforts. Although previous research has developed various models to predict greenhouse gas emissions, few studies have compared statistical models, machine learning, and deep learning techniques when applied to small sample datasets, particularly in the case of Taiwan, which typically has limited sample sizes. Accordingly, this study uses Taiwan’s small dataset to train predictive models, comparing statistical, machine learning, and deep learning approaches. This study significantly improved the models’ learning efficiency and predictive accuracy despite limited data resources by utilizing multiple feature selection techniques to identify highly predictive features and integrating extensive data preprocessing and data augmentation techniques. Results indicate that extreme gradient boosting outperforms other models, particularly under Gaussian noise conditions, achieving higher accuracy in mean squared error and mean absolute percentage error metrics. Although long short-term memory proves suitable for time series, it exhibits greater sensitivity to noise. Conversely, the non-linear Grey Bernoulli model (NGBM (1, N)) exhibits robust resistance to noise, although it demonstrates relatively lower predictive accuracy under standard conditions. This model aids Taiwan in tracking emission trends and informing climate policies, contributing valuable data insights for future environmental strategies. Moreover, this model demonstrates significant potential for application in environmental monitoring and climate policy formulation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Apply data science and feature selection techniques to predict carbon dioxide emissions in Taiwan

  • Shih-Hsien Tseng,
  • Pei Han Feng,
  • Thi Ha Trang Duong

摘要

Taiwan has introduced specific policies, such as the 2050 net-zero goal and the 2024 carbon fee, to support global climate action and encourage carbon reduction efforts. Although previous research has developed various models to predict greenhouse gas emissions, few studies have compared statistical models, machine learning, and deep learning techniques when applied to small sample datasets, particularly in the case of Taiwan, which typically has limited sample sizes. Accordingly, this study uses Taiwan’s small dataset to train predictive models, comparing statistical, machine learning, and deep learning approaches. This study significantly improved the models’ learning efficiency and predictive accuracy despite limited data resources by utilizing multiple feature selection techniques to identify highly predictive features and integrating extensive data preprocessing and data augmentation techniques. Results indicate that extreme gradient boosting outperforms other models, particularly under Gaussian noise conditions, achieving higher accuracy in mean squared error and mean absolute percentage error metrics. Although long short-term memory proves suitable for time series, it exhibits greater sensitivity to noise. Conversely, the non-linear Grey Bernoulli model (NGBM (1, N)) exhibits robust resistance to noise, although it demonstrates relatively lower predictive accuracy under standard conditions. This model aids Taiwan in tracking emission trends and informing climate policies, contributing valuable data insights for future environmental strategies. Moreover, this model demonstrates significant potential for application in environmental monitoring and climate policy formulation.