Exploring Interpretability in Backdoor Attacks on Image Classification
摘要
This study delves into the issue of backdoor attacks on deep learning models in image recognition, proposing defense strategies based on interpretability techniques. Through experiments and analysis, we verify the impacts of different types of backdoor attacks on deep neural networks, and explore the application effectiveness of gradient-based activation mapping techniques in enhancing model interpretability. This paper not only provides a deep theoretical foundation for understanding and detecting backdoor attacks but also offers practical defense approaches and methods for developing more robust and secure deep learning models. Future research can further explore complex attack scenarios and defense mechanisms to address evolving security challenges.