The classification of abusive and non-abusive speech in Dravidian languages presents a significant challenge, especially given the limited amount of labeled datasets and the complexity of the language. While there have been numerous studies focusing on abusive language detection in textual data, research in the audio domain remains limited. In this paper, we address abusive and non-abusive speech classification in Telugu language, which is largely unexplored. The novelty of this work lies in being the first to classify abusive speech in Telugu by introducing a novel dataset consisting of one hour each of abusive and non-abusive Telugu speech, comprising 221 and 232 audio files, respectively. Leveraging Vision Transformer for feature extraction and multi-layer perceptron for classification, we explore the classification of abusive and non-abusive speech in Telugu. Our proposed approach achieves a classification accuracy of up to 87.09% for the custom dataset. Our work is driven by the need to effectively detect harmful content in audio, contributing to safer online environments in alignment with content moderation policies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Abusive Speech Detection in Telugu Using Vision Transformers

  • A. Venkata Sai Krishna Varun,
  • Katragadda Geetha Deepika,
  • A. S. M. Dinakar,
  • Gunnam Himamsh,
  • G. Jyothish Lal

摘要

The classification of abusive and non-abusive speech in Dravidian languages presents a significant challenge, especially given the limited amount of labeled datasets and the complexity of the language. While there have been numerous studies focusing on abusive language detection in textual data, research in the audio domain remains limited. In this paper, we address abusive and non-abusive speech classification in Telugu language, which is largely unexplored. The novelty of this work lies in being the first to classify abusive speech in Telugu by introducing a novel dataset consisting of one hour each of abusive and non-abusive Telugu speech, comprising 221 and 232 audio files, respectively. Leveraging Vision Transformer for feature extraction and multi-layer perceptron for classification, we explore the classification of abusive and non-abusive speech in Telugu. Our proposed approach achieves a classification accuracy of up to 87.09% for the custom dataset. Our work is driven by the need to effectively detect harmful content in audio, contributing to safer online environments in alignment with content moderation policies.