Data leakage poses a critical threat to data security and privacy, necessitating robust detection mechanisms that safeguard sensitive information without compromising performance. This study evaluates the effectiveness of machine learning (ML) models in detecting data leakage using both AES-encrypted and non-encrypted datasets. The analysis covers traditional classifiers—Random Forest (RF), Support Vector Machine (SVM), and Decision Tree (DT)—alongside advanced models such as Gradient Boosting (GB) and XGBoost (XGB). Experimental results indicate minimal performance degradation due to encryption, with RF achieving \(62.54\%\) accuracy on encrypted data, closely matching its \(62.37\%\) accuracy on non-encrypted data. XGB outperformed other models, attaining \(71.22\%\) accuracy, although a recall of \(24.42\%\) resulted in a F1-score of \(34.55\%\) , highlighting a trade-off between precision and recall. The computational overhead introduced by encryption was also analyzed, showing that SVM experienced the highest overhead ( \(183.33\%\) ), while RF maintained a moderate increase ( \(33.33\%\) ). To mitigate these effects, Principal Component Analysis (PCA) was applied, balancing computational efficiency and predictive performance. The comprehensive evaluation confirms the feasibility of deploying ML models for secure data leakage detection with minimal performance compromise. These findings lay the groundwork for future research, including homomorphic encryption, federated learning, and deep learning architectures, to further enhance privacy-preserving capabilities. The study offers actionable insights for developing robust, scalable, and secure data leakage detection systems, essential for protecting sensitive data in the face of growing cybersecurity threats.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Privacy-Conscious Data Leakage Detection: An Experimental Evaluation of Machine Learning Models on Encrypted Data

  • Ahmed Abdallah Omer,
  • Akram Pasha

摘要

Data leakage poses a critical threat to data security and privacy, necessitating robust detection mechanisms that safeguard sensitive information without compromising performance. This study evaluates the effectiveness of machine learning (ML) models in detecting data leakage using both AES-encrypted and non-encrypted datasets. The analysis covers traditional classifiers—Random Forest (RF), Support Vector Machine (SVM), and Decision Tree (DT)—alongside advanced models such as Gradient Boosting (GB) and XGBoost (XGB). Experimental results indicate minimal performance degradation due to encryption, with RF achieving \(62.54\%\) accuracy on encrypted data, closely matching its \(62.37\%\) accuracy on non-encrypted data. XGB outperformed other models, attaining \(71.22\%\) accuracy, although a recall of \(24.42\%\) resulted in a F1-score of \(34.55\%\) , highlighting a trade-off between precision and recall. The computational overhead introduced by encryption was also analyzed, showing that SVM experienced the highest overhead ( \(183.33\%\) ), while RF maintained a moderate increase ( \(33.33\%\) ). To mitigate these effects, Principal Component Analysis (PCA) was applied, balancing computational efficiency and predictive performance. The comprehensive evaluation confirms the feasibility of deploying ML models for secure data leakage detection with minimal performance compromise. These findings lay the groundwork for future research, including homomorphic encryption, federated learning, and deep learning architectures, to further enhance privacy-preserving capabilities. The study offers actionable insights for developing robust, scalable, and secure data leakage detection systems, essential for protecting sensitive data in the face of growing cybersecurity threats.