Stochastic PAC-Bayesian transformers for network intrusion detection and natural language processing applications
摘要
Transformer architectures dominate contemporary machine learning, yet face critical limitations in security applications: vulnerability to adversarial attacks, lack of calibrated uncertainty estimates, and difficulty distinguishing confident predictions from uncertain cases requiring human review. We introduce stochastic probably approximately correct (PAC) Bayesian transformers that convert deterministic attention into probabilistic variants via variational inference, unifying uncertainty quantification, and adversarial robustness within a single framework. Our approach replaces fixed attention weights with learned variational distributions and propagates uncertainty through Monte Carlo (MC) sampling, creating moving targets that degrade adversarial effectiveness. We derive joint PAC-Bayesian bounds showing that parameter stochasticity improves both calibration and robustness, with complexity scaling as