The selection of optimal bandwidth \(\textit{bw}\) in Kernel Density Estimation (KDE) is crucial for accurate density estimation. Traditional methods, such as unbiased and biased cross-validation and their variants, often suffer from performance and computational limitations. This study presents a novel, data-driven approach to bandwidth selection using unsupervised machine learning techniques. Synthetic samples are generated from the original data, a comprehensive set of statistical parameters is calculated, and bandwidth is estimated using a rule-of-thumb [2]. Feature importance is assessed through random forests (RF) to identify the most influential factors in bandwidth selection. Various unsupervised learning models are then applied to predict the optimal bandwidth for the original sample \(\textit{bw}_{\text {os}}\) . The results demonstrate that this machine learning-based approach achieves superior accuracy and computational efficiency compared to traditional cross-validation methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data-Driven Bandwidth Selection in Kernel Density Estimation via Unsupervised Learning

  • Arsalane Chouaib Guidoum,
  • Amel Henider

摘要

The selection of optimal bandwidth \(\textit{bw}\) in Kernel Density Estimation (KDE) is crucial for accurate density estimation. Traditional methods, such as unbiased and biased cross-validation and their variants, often suffer from performance and computational limitations. This study presents a novel, data-driven approach to bandwidth selection using unsupervised machine learning techniques. Synthetic samples are generated from the original data, a comprehensive set of statistical parameters is calculated, and bandwidth is estimated using a rule-of-thumb [2]. Feature importance is assessed through random forests (RF) to identify the most influential factors in bandwidth selection. Various unsupervised learning models are then applied to predict the optimal bandwidth for the original sample \(\textit{bw}_{\text {os}}\) . The results demonstrate that this machine learning-based approach achieves superior accuracy and computational efficiency compared to traditional cross-validation methods.