Warning: This paper contains examples of the language that some people may find offensive. Social media has become an open platform for all its users. In the comment section, anyone can express their opinion, anger, frustration, and taunt toward a person or a group based on their religion, race, sex, or sexual orientation. Detecting this type of hate comment, which provokes people to become fierce unnecessarily, is another challenging target for all social media platforms like Facebook, Twitter, YouTube, etc. A little experiment is done in Indian languages, where European languages took massive advantage of being the most popular language in the world. The Assamese language is one of the languages in which hate speech detection work has not been done yet. This paper tried to explore Assamese hate, its issues, and challenges and finally perform cross-lingual testing on it. Firstly, we trained the state-of-the-art BERT (bidirectional encoder representations from transformers) models on the existing Bengali hate speech dataset and then evaluated the model’s performance on both Bengali and Assamese hate data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hate Speech Detection in the Assamese Language with Cross-Lingual Testing

  • Koyel Ghosh,
  • Apurbalal Senapati

摘要

Warning: This paper contains examples of the language that some people may find offensive. Social media has become an open platform for all its users. In the comment section, anyone can express their opinion, anger, frustration, and taunt toward a person or a group based on their religion, race, sex, or sexual orientation. Detecting this type of hate comment, which provokes people to become fierce unnecessarily, is another challenging target for all social media platforms like Facebook, Twitter, YouTube, etc. A little experiment is done in Indian languages, where European languages took massive advantage of being the most popular language in the world. The Assamese language is one of the languages in which hate speech detection work has not been done yet. This paper tried to explore Assamese hate, its issues, and challenges and finally perform cross-lingual testing on it. Firstly, we trained the state-of-the-art BERT (bidirectional encoder representations from transformers) models on the existing Bengali hate speech dataset and then evaluated the model’s performance on both Bengali and Assamese hate data.