The proliferation of fraudulent profiles on social media platforms presents noteworthy risks to security and erodes user confidence. The effectiveness of Random Forest, Support Vector Machine (SVM), and Neural Network models is the main emphasis of this study’s implementation and review of machine learning techniques for identifying bogus profiles. We prepared features, including followers, statuses, and account age, using data preprocessing, including encoding and normalization, on a Kaggle dataset. Our results show that, despite interpretable-city-city issues, the Random Forest classifier managed both continuous and categorical variables with an accuracy of 91%. With an RBF kernel, the SVM beat Random Forest with 93% accuracy, reducing false positives but consuming a large amount of processing power while optimizing hyperparameters. With its peak accuracy of 94%, the Neural Network proved to be an excellent tool for identifying intricate, non-linear patterns linked to fraudulent profiles. However, its interpretability and resource requirements were limited. A comparison investigation revealed that while the Neural Network provided more precision, Random Forest was more interpretable and performed better overall, making it appropriate for applications with lower processing capacity. While SVM is costly in computing resources, it is still favorably positioned due to its enhanced precision in minimizing false positives. This work highlights the challenges of feature selection in the context of the remaining difficulties of dataset imbalances and model explainability. Ultimately, a choice of a model rests on specific application requirements regarding, among other things, the clarity of decisions made, the economy of computation, and the precision of results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning for Fake Profile Users Detection in Social Network Systems: A Review and Implementation Phase

  • Aishwarya Waghmare,
  • Vinothkumar Kolluru,
  • Yagnesh Challagundla,
  • l. V. S. Aditya Bhrugumalla,
  • Advaitha Naidu Chintakunta,
  • Sagar Pande

摘要

The proliferation of fraudulent profiles on social media platforms presents noteworthy risks to security and erodes user confidence. The effectiveness of Random Forest, Support Vector Machine (SVM), and Neural Network models is the main emphasis of this study’s implementation and review of machine learning techniques for identifying bogus profiles. We prepared features, including followers, statuses, and account age, using data preprocessing, including encoding and normalization, on a Kaggle dataset. Our results show that, despite interpretable-city-city issues, the Random Forest classifier managed both continuous and categorical variables with an accuracy of 91%. With an RBF kernel, the SVM beat Random Forest with 93% accuracy, reducing false positives but consuming a large amount of processing power while optimizing hyperparameters. With its peak accuracy of 94%, the Neural Network proved to be an excellent tool for identifying intricate, non-linear patterns linked to fraudulent profiles. However, its interpretability and resource requirements were limited. A comparison investigation revealed that while the Neural Network provided more precision, Random Forest was more interpretable and performed better overall, making it appropriate for applications with lower processing capacity. While SVM is costly in computing resources, it is still favorably positioned due to its enhanced precision in minimizing false positives. This work highlights the challenges of feature selection in the context of the remaining difficulties of dataset imbalances and model explainability. Ultimately, a choice of a model rests on specific application requirements regarding, among other things, the clarity of decisions made, the economy of computation, and the precision of results.