In this paper, This study propose a novel approach to improve the safety of classification models by incorporating adversarial training into the model training process. This study demonstrate that adversarial training enhances the safety of models without significantly affecting their accuracy. The proposed approach methodology involves generating adversarial examples from the original dataset and creating a mixed training dataset that includes both original and adversarial examples. By evaluating the model’s performance on both types of datasets, This study can achieve a balance between accuracy and safety. The proposed approach offers a more resource-efficient alternative to traditional model tuning processes, which often require multiple iterations to meet performance standards. The proposed approach results indicate that including at least 20% adversarial examples in the training data effectively improves safety while allowing adjustments based on the relative importance of accuracy and safety. The new method simplifies the process of achieving safety requirements and reduces the resources needed for training and evaluation. Future research will focus on validating this approach with diverse safety-related datasets, various model structures, and different adversarial attack methods, as well as exploring its applicability to other non-functional requirements in AI.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Balancing Accuracy and Safety in AI: A Novel Adversarial Training Approach

  • Veera Venkata Naga Krishna Chaitanya Atkuri

摘要

In this paper, This study propose a novel approach to improve the safety of classification models by incorporating adversarial training into the model training process. This study demonstrate that adversarial training enhances the safety of models without significantly affecting their accuracy. The proposed approach methodology involves generating adversarial examples from the original dataset and creating a mixed training dataset that includes both original and adversarial examples. By evaluating the model’s performance on both types of datasets, This study can achieve a balance between accuracy and safety. The proposed approach offers a more resource-efficient alternative to traditional model tuning processes, which often require multiple iterations to meet performance standards. The proposed approach results indicate that including at least 20% adversarial examples in the training data effectively improves safety while allowing adjustments based on the relative importance of accuracy and safety. The new method simplifies the process of achieving safety requirements and reduces the resources needed for training and evaluation. Future research will focus on validating this approach with diverse safety-related datasets, various model structures, and different adversarial attack methods, as well as exploring its applicability to other non-functional requirements in AI.