One of the most delicate concerns is cyberbullying due to today’s worldwide web advancement. Bullying may contribute to severe consequences, including mental health issues, academic performance challenges, and job dropout. It can also instill hatred in individuals or communities, making cyberbullying prevention critical. Our main goal in this paper is to mitigate the impact of cyberbullying and contribute to fostering a healthier Internet environment where individuals can interact without fear of unwarranted criticism. We propose a proactive approach involving early detection mechanisms. In this study, we employ a comprehensive analysis of tweets using multiple factors. The data is preprocessed using machine learning methods and natural language processing (NLP), followed by Term Frequency Inverse Document Frequency (TF-IDF) and Count Vectorizer. Several machine learning models, including Light Gradient Boosting Machine (LightGBM) and 12 others, are evaluated. The dataset comprises over 47,000 tweets from various perspectives. Our analysis indicates that the LightGBM classifier outperforms other models with an accuracy of 94.24%, precision of 95%, recall of 94%, and an F1-score of 94%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncovering Bullying on Social Media Platforms: A Comprehensive Study of Machine Learning Classifiers for Cyberbullying Detection

  • Hasibul Hamim,
  • Khandaker Mohammad Mohi Uddin,
  • Mst. Nishat Tasnim Mim,
  • Rafid Mostafiz,
  • Md. Abdul Based

摘要

One of the most delicate concerns is cyberbullying due to today’s worldwide web advancement. Bullying may contribute to severe consequences, including mental health issues, academic performance challenges, and job dropout. It can also instill hatred in individuals or communities, making cyberbullying prevention critical. Our main goal in this paper is to mitigate the impact of cyberbullying and contribute to fostering a healthier Internet environment where individuals can interact without fear of unwarranted criticism. We propose a proactive approach involving early detection mechanisms. In this study, we employ a comprehensive analysis of tweets using multiple factors. The data is preprocessed using machine learning methods and natural language processing (NLP), followed by Term Frequency Inverse Document Frequency (TF-IDF) and Count Vectorizer. Several machine learning models, including Light Gradient Boosting Machine (LightGBM) and 12 others, are evaluated. The dataset comprises over 47,000 tweets from various perspectives. Our analysis indicates that the LightGBM classifier outperforms other models with an accuracy of 94.24%, precision of 95%, recall of 94%, and an F1-score of 94%.