Abusive Speech Detection in Telugu Using Vision Transformers
摘要
The classification of abusive and non-abusive speech in Dravidian languages presents a significant challenge, especially given the limited amount of labeled datasets and the complexity of the language. While there have been numerous studies focusing on abusive language detection in textual data, research in the audio domain remains limited. In this paper, we address abusive and non-abusive speech classification in Telugu language, which is largely unexplored. The novelty of this work lies in being the first to classify abusive speech in Telugu by introducing a novel dataset consisting of one hour each of abusive and non-abusive Telugu speech, comprising 221 and 232 audio files, respectively. Leveraging Vision Transformer for feature extraction and multi-layer perceptron for classification, we explore the classification of abusive and non-abusive speech in Telugu. Our proposed approach achieves a classification accuracy of up to 87.09% for the custom dataset. Our work is driven by the need to effectively detect harmful content in audio, contributing to safer online environments in alignment with content moderation policies.