Implementation of Hybrid Machine Learning Algorithms in Classification of Real and Fake Profiles
摘要
Social media platforms have been increasingly popular in recent times, with millions of people logging on daily to connect with one another. But there has also been an increase in the quantity of fake accounts in recent times, particularly during the pandemic when everyone was online, leaving them exposed to various attacks from bots, malicious users, and other fake accounts. They can be used for a wide range of serious undesirable activities, such as propagating false information, engaging in fraud, cyberbullying, upsetting democratic procedures, and endorsing extreme viewpoints. Differentiating between these types of accounts and the real ones has become critical. This research work seeks to define characteristics that distinguish authentic accounts from fraudulent ones and also to implement hybrid algorithms for improved performance and better distinguishing the genuineness of the profiles. This research uses symmetric uncertainty feature selection approach to identify attributes that distinguish a real account from a fake account. The performance of several legacy algorithms, including Random Forest, SVM, Decision Tree, Neural Network, and Logistic Regression, is compared after implementation. The most effective combination of these algorithms is then discovered by combining them using Voting Classifier. It has been found that the best hybrid model for effectively classifying fake and real profiles was a combination of Random Forest and Neural Network.