<p>In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional languages, leading to code-mixed text. While this linguistic flexibility fosters freedom of expression, it also contributes to the rampant spread of hate speech, posing significant societal challenges. Identifying hate speech in low-resource Kannada-English code-mixed text is challenging due to the scarcity of annotated corpora. To address this gap, the authors introduce HASTIKA (Hate Speech and Target Identification in Kannada-English Code-Mixed Text), a gold-standard corpus specifically designed for hate speech detection and target identification. It consists of 8,058 YouTube comments, annotated for binary classification “Hate" and “Non-hate" and fine-grained categorization into “Gender", “Political", “Religion", “Geo-political", “Violence", and “Others", marking a significant contribution as the first dataset exclusively tailored for this purpose. Topic modeling techniques were applied to uncover latent themes. Benchmark experiments show that fastText, which captures word-level features, achieves 0.7529 accuracy for binary classification and 0.6042 for multi-class classification, while BERT, excelling in sentence-level feature extraction, attains 0.8054 and 0.6819 accuracy, respectively. This research provides a foundational resource for hate speech detection in Kannada-English code-mixed text and underscores the importance of linguistic and contextual information in developing robust classification models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HASTIKA: hate speech and target identification in Kannada-English code-mixed text

  • Sanjana Kavatagi,
  • Rashmi Rachh

摘要

In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional languages, leading to code-mixed text. While this linguistic flexibility fosters freedom of expression, it also contributes to the rampant spread of hate speech, posing significant societal challenges. Identifying hate speech in low-resource Kannada-English code-mixed text is challenging due to the scarcity of annotated corpora. To address this gap, the authors introduce HASTIKA (Hate Speech and Target Identification in Kannada-English Code-Mixed Text), a gold-standard corpus specifically designed for hate speech detection and target identification. It consists of 8,058 YouTube comments, annotated for binary classification “Hate" and “Non-hate" and fine-grained categorization into “Gender", “Political", “Religion", “Geo-political", “Violence", and “Others", marking a significant contribution as the first dataset exclusively tailored for this purpose. Topic modeling techniques were applied to uncover latent themes. Benchmark experiments show that fastText, which captures word-level features, achieves 0.7529 accuracy for binary classification and 0.6042 for multi-class classification, while BERT, excelling in sentence-level feature extraction, attains 0.8054 and 0.6819 accuracy, respectively. This research provides a foundational resource for hate speech detection in Kannada-English code-mixed text and underscores the importance of linguistic and contextual information in developing robust classification models.