<p>Hamiltonian Monte Carlo [HMC] is widely regarded as the de facto standard for posterior sampling-based inference [SBI] of Bayesian neural networks [BNNs]. Iterative gradient computations required in HMC to generate proposals can become prohibitively expensive, particularly for BNNs with numerous latent nodes and thus high-dimensional posterior distributions. We consider a gradient-free approach, leveraging a generalized auxiliary model that admits tractable full conditional distributions to devise a Metropolis-within-Gibbs sampler with asymmetric proposals that do not require gradients. Through simulation studies with shallow-wide BNNs, we demonstrate that the proposed sampler produces posterior samples of superior or comparable quality to HMC in key aspects such as nonlinearity approximation, out-of-distribution uncertainty quantification, acceptance rates, and effective sample sizes, while being gradient-free and thus significantly cost-efficient, making it a practical alternative for BNN posterior inference.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Gradient-free Gibbs sampler for shallow-wide Bayesian neural networks

  • Geonhee Han

摘要

Hamiltonian Monte Carlo [HMC] is widely regarded as the de facto standard for posterior sampling-based inference [SBI] of Bayesian neural networks [BNNs]. Iterative gradient computations required in HMC to generate proposals can become prohibitively expensive, particularly for BNNs with numerous latent nodes and thus high-dimensional posterior distributions. We consider a gradient-free approach, leveraging a generalized auxiliary model that admits tractable full conditional distributions to devise a Metropolis-within-Gibbs sampler with asymmetric proposals that do not require gradients. Through simulation studies with shallow-wide BNNs, we demonstrate that the proposed sampler produces posterior samples of superior or comparable quality to HMC in key aspects such as nonlinearity approximation, out-of-distribution uncertainty quantification, acceptance rates, and effective sample sizes, while being gradient-free and thus significantly cost-efficient, making it a practical alternative for BNN posterior inference.