<p>This study aims to classify natural and non-natural seismic waveforms under imbalanced conditions. We will optimize oversampling and downsampling rates using the Sparrow Search Algorithm (SSA) and ADASYN data augmentation. These methods will be integrated with tree-based classification models to improve accuracy. Through the research in this article, it was found that the combination of SSA–ADASYN–XGBoost and the permutation entropy extraction technique with a sliding window of 500 can achieve a seismic event binary classification accuracy of 0.97. SSA performs better than optimization models such as Particle Swarm Optimization (PSO) and Genetic Algorithm (GA), ADASYN adaptive data augmentation method performs better than SMOTE and other methods in seismic classification, while XGBoost and LightGBM have higher accuracy than random forest, and XGBoost uses less proportion of data augmentation compared to LightGBM, which to some extent reduces the possibility of overfitting.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Seismic source attribute recognition with signal processing and tree-based machine learning on imbalanced samples

  • Wei Chen,
  • Xinlong Zhang,
  • Qi Shao

摘要

This study aims to classify natural and non-natural seismic waveforms under imbalanced conditions. We will optimize oversampling and downsampling rates using the Sparrow Search Algorithm (SSA) and ADASYN data augmentation. These methods will be integrated with tree-based classification models to improve accuracy. Through the research in this article, it was found that the combination of SSA–ADASYN–XGBoost and the permutation entropy extraction technique with a sliding window of 500 can achieve a seismic event binary classification accuracy of 0.97. SSA performs better than optimization models such as Particle Swarm Optimization (PSO) and Genetic Algorithm (GA), ADASYN adaptive data augmentation method performs better than SMOTE and other methods in seismic classification, while XGBoost and LightGBM have higher accuracy than random forest, and XGBoost uses less proportion of data augmentation compared to LightGBM, which to some extent reduces the possibility of overfitting.