Weighted Naive Bayes methods have recently been developed to alleviate the strong conditional independence assumption of traditional Naive Bayes classifiers. In particular, class-specific attribute weighted Naive Bayes (CAWNB) has been shown to yield excellent performance on many modern datasets. Such methods, however, are prone to over-fitting on small sample, large feature space data. In this work, we propose a Bayesian Regularized Iterative Shrinkage-Thresholding Algorithm (BARISTA), which includes both \(\ell _1\) and \(\ell _2\) regularization to mitigate this problem. As we show, estimating the parameters of BARISTA via maximum likelihood yields a convex objective that can be efficiently optimized using Iterative Shrinkage-Thresholding Algorithms (ISTA). We prove the resulting method has many attractive theoretical and numerical properties, including a guaranteed linear rate of convergence. Using several standard benchmark datasets, we demonstrate how BARISTA can yield a significant increase in performance compared to many state-of-the-art weighted Naive Bayes methods. We also show how the Fast Iterative-Shrinkage Thresholding Algorithm (FISTA) can be used to further accelerate convergence. (Our code and data are publicly available on this repository.)

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bayesian Regularized Iterative Soft Thresholding Algorithm

  • Nicolas Cutrona,
  • Dominique Guillot

摘要

Weighted Naive Bayes methods have recently been developed to alleviate the strong conditional independence assumption of traditional Naive Bayes classifiers. In particular, class-specific attribute weighted Naive Bayes (CAWNB) has been shown to yield excellent performance on many modern datasets. Such methods, however, are prone to over-fitting on small sample, large feature space data. In this work, we propose a Bayesian Regularized Iterative Shrinkage-Thresholding Algorithm (BARISTA), which includes both \(\ell _1\) and \(\ell _2\) regularization to mitigate this problem. As we show, estimating the parameters of BARISTA via maximum likelihood yields a convex objective that can be efficiently optimized using Iterative Shrinkage-Thresholding Algorithms (ISTA). We prove the resulting method has many attractive theoretical and numerical properties, including a guaranteed linear rate of convergence. Using several standard benchmark datasets, we demonstrate how BARISTA can yield a significant increase in performance compared to many state-of-the-art weighted Naive Bayes methods. We also show how the Fast Iterative-Shrinkage Thresholding Algorithm (FISTA) can be used to further accelerate convergence. (Our code and data are publicly available on this repository.)