Local Intrinsic Dimensionality (LID) is a measure of data complexity in the vicinity of a query point. In this work, we address the problem of estimating LID from a Bayesian perspective by establishing a theoretical framework that derives the distribution of LID given a data sample. Using this framework, we develop new LID estimators that can outperform the Maximum Likelihood Estimator (MLE) in certain contexts. The framework also provides a convenient way to incorporate prior LID knowledge through informative priors. Additionally, we demonstrate how to aggregate multiple LID distributions in a Bayesian manner using logarithmic pooling. We conduct a variety of experiments, demonstrating that a Bayesian approach to LID is effective with a small number of nearest neighbors and when incorporating informative priors. We also show that in deep neural networks, MLE produces highly volatile LID estimates, whereas a Bayesian approach that incorporates prior LID information smoothes and reduces the variance of these estimates.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bayesian Estimation Approaches for Local Intrinsic Dimensionality

  • Zaher Joukhadar,
  • Hanxun Huang,
  • Sarah Monazam Erfani,
  • Ricardo J. G. B. Campello,
  • Michael E. Houle,
  • James Bailey

摘要

Local Intrinsic Dimensionality (LID) is a measure of data complexity in the vicinity of a query point. In this work, we address the problem of estimating LID from a Bayesian perspective by establishing a theoretical framework that derives the distribution of LID given a data sample. Using this framework, we develop new LID estimators that can outperform the Maximum Likelihood Estimator (MLE) in certain contexts. The framework also provides a convenient way to incorporate prior LID knowledge through informative priors. Additionally, we demonstrate how to aggregate multiple LID distributions in a Bayesian manner using logarithmic pooling. We conduct a variety of experiments, demonstrating that a Bayesian approach to LID is effective with a small number of nearest neighbors and when incorporating informative priors. We also show that in deep neural networks, MLE produces highly volatile LID estimates, whereas a Bayesian approach that incorporates prior LID information smoothes and reduces the variance of these estimates.