The Internet and social media have made significant progress and are now accessible to a wide audience. With this rapid growth there has also been a corresponding increase in negativity to a certain degree. The main goal is to eliminate the issue of Cyberbullying, which often occurs via the use of Offensive language. Although several options exist, many of them are currently under development. Nevertheless, the issue is always changing, particularly for low-resource languages like Bengali. It is crucial to discover more advanced and enhanced solutions since there are always constraints with present methods and opportunities for further enhancement. This work presents a novel approach to detecting foul language by using a combination of two models in a multi-step solution. Applying two distinct classifiers enables a more detailed method for identifying offensive language. Between the two models, one may ascertain the character of the statement by categorising it as either offensive or non-offensive. The second approach employs multi-label classification to categorise objectionable words into four distinct categories: “Sexual”, “Dishonour”, “Incursive”, and “Neutral” hence advancing the classification process. Several models were used in the process, while Recurrent Neural Networks (RNN) exhibited superior outcomes and were used for both models. The accuracy for both the offensive and non-offensive models was 87.37%. In terms of Multi-label classification, the accuracy achieved was 67%. Although the performance of the multi-label classification model was comparitivly low, however the combination of these two models demonstrates remarkable performance in terms detecting Offensive Bengali language.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Refining Bengali Hate Speech Detection: Multi-label Classification Using RNN and LSTM

  • Mominul Islam,
  • Md Sanjid Hossain,
  • Md Monjurul Islam,
  • Md Abul Hossain Murad,
  • Sadman Sakib Niloy,
  • Md. Sanzidul Islam,
  • Touhid Bhuiyan Islam

摘要

The Internet and social media have made significant progress and are now accessible to a wide audience. With this rapid growth there has also been a corresponding increase in negativity to a certain degree. The main goal is to eliminate the issue of Cyberbullying, which often occurs via the use of Offensive language. Although several options exist, many of them are currently under development. Nevertheless, the issue is always changing, particularly for low-resource languages like Bengali. It is crucial to discover more advanced and enhanced solutions since there are always constraints with present methods and opportunities for further enhancement. This work presents a novel approach to detecting foul language by using a combination of two models in a multi-step solution. Applying two distinct classifiers enables a more detailed method for identifying offensive language. Between the two models, one may ascertain the character of the statement by categorising it as either offensive or non-offensive. The second approach employs multi-label classification to categorise objectionable words into four distinct categories: “Sexual”, “Dishonour”, “Incursive”, and “Neutral” hence advancing the classification process. Several models were used in the process, while Recurrent Neural Networks (RNN) exhibited superior outcomes and were used for both models. The accuracy for both the offensive and non-offensive models was 87.37%. In terms of Multi-label classification, the accuracy achieved was 67%. Although the performance of the multi-label classification model was comparitivly low, however the combination of these two models demonstrates remarkable performance in terms detecting Offensive Bengali language.