The recent AI advancements, most notably Foundation Models, were naturally followed by an increased demand of AI-based solutions, which became widely available. However, the entities who offer these models are exposed to the risk of their models being stolen. Knowledge distillation demonstrated great benefits in reducing the inference time of deep neural networks, hence it has been an area of great interest. Model stealing represents a sub-category of knowledge distillation, which has a malign purpose: extracting the capability of a black-box model (teacher) into a local model (student). In this context, we propose a framework for stealing a model (which is accessible through an API) under very strict constraints (e.g.  no access to the original training data, the architecture or the weights of the model), with a focus on making as few calls as possible to the teacher model. We first generate synthetic data using diffusion models, a powerful class of generative models showcasing strong capabilities in image synthesis. Although they have been extensively applied to many tasks in computer vision, we propose to explore a new use case in our work, namely generating an artificial data set (called proxy data set) for knowledge distillation. Next, we obtain the labels for synthetic samples by passing them through the black-box model. More precisely, we only collect the predictions (either soft or hard) for a number image samples, as we have to comply with a fixed number of allowed API calls. In our framework, we introduce a novel active self-paced learning mechanism to thoroughly utilize the proxy data during distillation to its complete potential. The final step consists of distilling the knowledge of the black-box teacher (attacked model) into a student model (copy of the attacked model). Besides the labeled collected data, we also utilize the remaining unlabeled data generated by the diffusion model. Our empirical results on two data sets with different characteristics confirm the superiority of our framework over two state-of-the-art methods in the few-call model extraction scenario.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Active Self-paced Knowledge Distillation and Diffusion-Based Generation for Few-Call Model Stealing

  • Vlad Hondru,
  • Radu-Tudor Ionescu

摘要

The recent AI advancements, most notably Foundation Models, were naturally followed by an increased demand of AI-based solutions, which became widely available. However, the entities who offer these models are exposed to the risk of their models being stolen. Knowledge distillation demonstrated great benefits in reducing the inference time of deep neural networks, hence it has been an area of great interest. Model stealing represents a sub-category of knowledge distillation, which has a malign purpose: extracting the capability of a black-box model (teacher) into a local model (student). In this context, we propose a framework for stealing a model (which is accessible through an API) under very strict constraints (e.g.  no access to the original training data, the architecture or the weights of the model), with a focus on making as few calls as possible to the teacher model. We first generate synthetic data using diffusion models, a powerful class of generative models showcasing strong capabilities in image synthesis. Although they have been extensively applied to many tasks in computer vision, we propose to explore a new use case in our work, namely generating an artificial data set (called proxy data set) for knowledge distillation. Next, we obtain the labels for synthetic samples by passing them through the black-box model. More precisely, we only collect the predictions (either soft or hard) for a number image samples, as we have to comply with a fixed number of allowed API calls. In our framework, we introduce a novel active self-paced learning mechanism to thoroughly utilize the proxy data during distillation to its complete potential. The final step consists of distilling the knowledge of the black-box teacher (attacked model) into a student model (copy of the attacked model). Besides the labeled collected data, we also utilize the remaining unlabeled data generated by the diffusion model. Our empirical results on two data sets with different characteristics confirm the superiority of our framework over two state-of-the-art methods in the few-call model extraction scenario.