ProvDP: Differential Privacy for System Provenance Dataset
摘要
Provenance-based Intrusion Detection System (PIDS) is getting widely deployed for safeguarding enterprises across various verticals against sophisticated cyberattacks such as Advanced Persistent Threat (APT) campaigns. Rich in detail, provenance data are essentially fine-grained activities of users logged at their endpoints. Therefore, a security breach of provenance data can result in the leak of private information pertaining to users and the enterprise they work for (e.g., the clients to which a user often communicates), raising significant privacy concerns. In this work, we propose a novel privacy-preserving solution, ProvDP, specifically tailored to protect the privacy of provenance data and focusing on provenance graphs. Our approach introduces multiple technical contributions: (1) a novel method for converting provenance graphs to and from provenance trees, enabling us to harness properties of trees to develop privacy-preserving techniques while preserving the semantic value of provenance graphs, (2) a novel subtree differential privacy framework for providing privacy guarantee on these provenance trees, and (3) empirical evidence that the application of differential privacy does not diminish the detection accuracy of PIDSes. Our evaluation demonstrates that PIDS trained on differentially private data maintain utility while preserving privacy.