Cyberbullying, characterized by digital abuse such as harassment and doxing, has become prevalent on social media platforms, mainly targeting despised groups. Victims often endure severe psychological effects, including anxiety and strained interpersonal relationships, sometimes ending in tragic outcomes like suicide. To mitigate these issues, automated systems for detecting cyberbullying text are crucial. While recent methods have employed classical, deep learning, and transformer-based language models like BERT, there remains a gap in the literature regarding the comparative effective-ness of large language models in this domain. This study addresses this gap by evaluating the efficacy of large language models, specifically Mistral 7B and Llama3, against the transformer-based model BERT. The comparison encompasses binary and multiclass classification scenarios, assessing their performance in identifying cyberbullying content. The multiclass BERT model has outperformed the literature's large language and other benchmark models, achieving an F1 score of 83.67%. The BERT model was capable of classifying multiple classes effectively without being biased.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Language Model-Based Approach for Multiclass Cyberbullying Detection

  • Sanaa Kaddoura,
  • Reem Nassar

摘要

Cyberbullying, characterized by digital abuse such as harassment and doxing, has become prevalent on social media platforms, mainly targeting despised groups. Victims often endure severe psychological effects, including anxiety and strained interpersonal relationships, sometimes ending in tragic outcomes like suicide. To mitigate these issues, automated systems for detecting cyberbullying text are crucial. While recent methods have employed classical, deep learning, and transformer-based language models like BERT, there remains a gap in the literature regarding the comparative effective-ness of large language models in this domain. This study addresses this gap by evaluating the efficacy of large language models, specifically Mistral 7B and Llama3, against the transformer-based model BERT. The comparison encompasses binary and multiclass classification scenarios, assessing their performance in identifying cyberbullying content. The multiclass BERT model has outperformed the literature's large language and other benchmark models, achieving an F1 score of 83.67%. The BERT model was capable of classifying multiple classes effectively without being biased.