<p>Non-technical losses (NTL) in power distribution, such as illegal meter tapping, cause significant financial losses for utilities, amounting to billions annually. This study evaluates various machine learning methods for NTL detection, addressing the challenge of imbalanced electricity consumption data. Seven techniques for data balancing were employed: Adaptive Synthetic Sampling (ADASYN), Random Over Sampling, Random Under Sampling, Near Miss Under Sampling, and several variations of Synthetic Minority Over Sampling (SMOTE), including Borderline-SMOTE, SMOTE-ENN, and SMOTE-Tomek links. The model comprises two stages: first, seven classification algorithms (Decision Tree, Logistic Regression, XGBoost, Random Forest, SVM, Naïve Bayes, and KNN) were tested across diverse training-testing ratios to identify optimal performance. The second stage applied the comprehensive consumption dataset along with data balancing techniques to improve algorithm efficacy. Performance metrics—accuracy, precision, recall, F1 score, and Matthews Correlation Coefficient (MCC)—were utilized for evaluation. Results revealed that the Random Forest algorithm, when paired with Random Over Sampling at a 70 − 30% training-testing ratio, yielded the highest metrics: 98.03% accuracy, 99.02% precision, surpassing existing literature. The model achieved exceptional precision (0.990) and the highest overall performance, with rigorous statistical testing confirming all improvements were significant at the 95% confidence level.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Non-technical loss detection in power distribution networks using machine learning

  • Safdar Ali Abro,
  • Javed Ahmed Laghari,
  • Sufyan Ali Memon,
  • Talha Ahmed Khan,
  • Imran Memon,
  • Haidawati Nasir,
  • Kaneez Fatima

摘要

Non-technical losses (NTL) in power distribution, such as illegal meter tapping, cause significant financial losses for utilities, amounting to billions annually. This study evaluates various machine learning methods for NTL detection, addressing the challenge of imbalanced electricity consumption data. Seven techniques for data balancing were employed: Adaptive Synthetic Sampling (ADASYN), Random Over Sampling, Random Under Sampling, Near Miss Under Sampling, and several variations of Synthetic Minority Over Sampling (SMOTE), including Borderline-SMOTE, SMOTE-ENN, and SMOTE-Tomek links. The model comprises two stages: first, seven classification algorithms (Decision Tree, Logistic Regression, XGBoost, Random Forest, SVM, Naïve Bayes, and KNN) were tested across diverse training-testing ratios to identify optimal performance. The second stage applied the comprehensive consumption dataset along with data balancing techniques to improve algorithm efficacy. Performance metrics—accuracy, precision, recall, F1 score, and Matthews Correlation Coefficient (MCC)—were utilized for evaluation. Results revealed that the Random Forest algorithm, when paired with Random Over Sampling at a 70 − 30% training-testing ratio, yielded the highest metrics: 98.03% accuracy, 99.02% precision, surpassing existing literature. The model achieved exceptional precision (0.990) and the highest overall performance, with rigorous statistical testing confirming all improvements were significant at the 95% confidence level.