According to the development in the field of social media, new generations have begun using several techniques in addition to written texts to express feelings, including hatred. Emojis are one of the most popular symbols that are used intensively to express feelings and emotions more clearly. All hate speech detection models generate word embeddings from the input text. Unfortunately, most of these models do not generate emoji embeddings. Instead, they either ignore or convert them to other forms. In this paper, we proposed a novel model, EMOJI-RoBERTa, to generate contextual embeddings for emojis beside words. These embeddings are learned by further pre-training the RoBERTa model using large emoji-rich unlabeled datasets and fine-tuning using labeled hate speech detection datasets. Results show that our model outperformed the conventional ones with accuracy and an F1 score of 92% each on the HASOC2020 dataset and 68% each on the HATEMOJI dataset. The system can be improved by using larger unlabeled and labeled social media datasets containing emojis for pre-training and fine-tuning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emoji-RoBERTa: A Contextual Emoji Representation for Hate Speech Detection

  • Jinan Ali Aljawazeri,
  • Mahdi Nsaif Jasim

摘要

According to the development in the field of social media, new generations have begun using several techniques in addition to written texts to express feelings, including hatred. Emojis are one of the most popular symbols that are used intensively to express feelings and emotions more clearly. All hate speech detection models generate word embeddings from the input text. Unfortunately, most of these models do not generate emoji embeddings. Instead, they either ignore or convert them to other forms. In this paper, we proposed a novel model, EMOJI-RoBERTa, to generate contextual embeddings for emojis beside words. These embeddings are learned by further pre-training the RoBERTa model using large emoji-rich unlabeled datasets and fine-tuning using labeled hate speech detection datasets. Results show that our model outperformed the conventional ones with accuracy and an F1 score of 92% each on the HASOC2020 dataset and 68% each on the HATEMOJI dataset. The system can be improved by using larger unlabeled and labeled social media datasets containing emojis for pre-training and fine-tuning.