Dimensionality Reduction in Environmental Data
摘要
Species distribution models (SDMs), often termed habitat suitability models, are extensively employed in ecology for diverse aims, including species conservation, habitat exploration, and the acquisition of evolutionary insights by estimating species distributions. These models shed light on the subject’s ecological and evolutionary dimensions. Feature selection (FS) seeks to lower model costs and its demand for storage, pick relevant interconnected features or eliminate superfluous and repetitive ones, as well as improve the resulting model’s clarity. Consequently, in order to estimate the range of seven bird species, this study identifies the best combination (classifier-filter) using five filter-based univariate feature selection techniques to choose relevant features with 40% and 50% thresholds as well as four classifiers: Random Forest (RF), Light gradient-boosting machine (LGBM), Decision Tree (DT), and Support Vector Machine (SVM). The empirical tests involve several methods, such as 5-fold cross-validation method, Scott Knott statistical (SK) test, and Borda Count voting method. In addition, we employed three performance metrics (accuracy, kappa, and F1-score). Experiments confirm that the filter Fisher-score gave satisfactory results regardless of the classifier used. In addition, Annual Mean Temperature, Mean Diurnal Range, Isothermally, Precipitation of Driest Month, and Land Cover were the most relevant features for predicting species distribution.