In recent years, there has been a significant increase in toxic and hateful speech on social media platforms, becoming deeply entrenched in online interactions. This issue has drawn the attention of researchers from various academic fields, leading them to extend their focus to include disciplines such as Natural Language Processing, Machine Learning, and Linguistics, in addition to traditional areas like Law, Sociology, Psychology, and Politics. This paper introduces an approach for detecting toxic and hateful speech on social media using Tabular Deep Learning. The goal is to apply and evaluate the performance of the FT-Transformer model in detecting hateful and toxic content in textual comments on social media in Brazilian Portuguese. An important aspect of this research involves using modern embedding models as language embedders and language models and evaluating their performance with the FT-Transformer, a transformer-based tabular model. The experimental scenario uses a binary version of the ToLD-Br dataset. Our approach achieved a 76% accuracy rate and a 75% macro F1-score using the OpenAI text-embedding-3-large model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Transformer-Based Tabular Approach to Detect Toxic Comments

  • Ghivvago Damas,
  • Rafael Torres Anchiêta,
  • Raimundo Santos Moura,
  • Vinicius Ponte Machado

摘要

In recent years, there has been a significant increase in toxic and hateful speech on social media platforms, becoming deeply entrenched in online interactions. This issue has drawn the attention of researchers from various academic fields, leading them to extend their focus to include disciplines such as Natural Language Processing, Machine Learning, and Linguistics, in addition to traditional areas like Law, Sociology, Psychology, and Politics. This paper introduces an approach for detecting toxic and hateful speech on social media using Tabular Deep Learning. The goal is to apply and evaluate the performance of the FT-Transformer model in detecting hateful and toxic content in textual comments on social media in Brazilian Portuguese. An important aspect of this research involves using modern embedding models as language embedders and language models and evaluating their performance with the FT-Transformer, a transformer-based tabular model. The experimental scenario uses a binary version of the ToLD-Br dataset. Our approach achieved a 76% accuracy rate and a 75% macro F1-score using the OpenAI text-embedding-3-large model.