Machine Unlearning for Foundation Models
摘要
This chapter aims to provide a comprehensive understanding of emerging machine unlearning (MU) techniques in foundation models. These techniques are designed to precisely evaluate the impact of specific data and high-level knowledge concepts on model performance and to efficiently and effectively eliminate their (possibly harmful) influence within a pre-trained model in response to users’ removal requests. Initially proposed to address data privacy concerns in compliance with the ‘right to be forgotten’ regulation, MU has become increasingly crucial with the advent of foundation models, as re-training from scratch (after removing the undesired training points) is prohibitively expensive in terms of time, compute, and money. In this section, we explore MU from several key perspectives: foundational concepts and formulations, optimization techniques, adversarial evaluation methods, and practical applications.